{
  "id": 667624,
  "title": "Design an optimal ground truth mask for ECG signal segmentation (30 dB barrier?)",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/667624",
  "author_name": "",
  "post_date": "2026-01-13T13:26:52.454502900Z",
  "votes": 6,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi everyone 👋,</p>\n<p>I wanted to share how I’m currently constructing pixel-level ground truth masks from the provided voltage CSVs, and also hear how others are thinking about designing targets for this task.</p>\n<p><strong>Converting voltage signals into pixel-level ground truth</strong></p>\n<p>The competition provides voltage time series per lead, but the models operate on images. So the first step is converting the voltage data into pixel coordinates aligned with the rendered ECG images.</p>\n<p>What I’m doing is:</p>\n<ul>\n<li>Converting each CSV into .dat and .hea files.</li>\n<li>Using ecg_image_kit to render the ECG and extract pixel coordinates.</li>\n</ul>\n<p>For example:</p>\n<pre><code>python gen_ecg_image_from_data.py \\\n  -i root/physionet-ecg-image-digitization/train/7663343/7663343.dat \\\n  -hea root/physionet-ecg-image-digitization/train/7663343/7663343.hea \\\n  -o /root/test \\\n  -st 0 \\\n  --print_header \\\n  --lead_bbox \\\n  --lead_name_bbox \\\n  --store_config 2 \\\n  --calibration_pulse 1\n</code></pre>\n<p>This produces a JSON file with plotted_pixels per lead, which gives the exact rasterized location of the ECG traces in the image.\n(Alternatively, it’s possible to reconstruct the pixel positions directly by following the rendering logic from ecg_image_kit: lead layout + scaling of ~78.74 pixels per mV at 200 DPI.)</p>\n<p><strong>Building the segmentation mask</strong></p>\n<p>From the extracted pixel coordinates, I generate a multi-channel mask:</p>\n<ul>\n<li>Resolution: 1700 × 2200 (same as the input image)</li>\n<li>Channels: 12 leads + long lead II + background</li>\n</ul>\n<p>I draw each lead trace as a thin line using either:\n-a 1-pixel binary line, or\n-an antialiased line (e.g. matplotlib with linewidth ≈ 0.75).</p>\n<p>Optionally, I upscale the width (e.g. ×2) before training to improve temporal resolution.</p>\n<p>This gives me a clean pixel-level target aligned with the ECG image geometry.</p>\n<p><strong>Training and extracting the signal</strong></p>\n<p>I train a standard segmentation network (U-Net style), then extract the signal per column using argmax, soft-argmax, or other methods. Finally I convert pixel positions back to voltage using the known scale (≈78.74 px = 1 mV), resample to the target frequency, and evaluate SNR. This pipeline is reasonably stable and gives a consistent baseline (around ~14 dB on the leaderboard for me so far).</p>\n<p><strong>A question for discussion</strong></p>\n<p>I’ve seen people mention ideas like “high-resolution targets”, “pixel nudging”, or “refining the labels with a network”, and I’d be very interested to hear how others are approaching this part. I have done some attempts but cannot go beyond 25-30dB barrier as ground-truth segmentation. What I’m especially curious about is how others are thinking about the design of the ground truth itself.</p>\n<p>Thanks for reading — and thanks in advance if you’re willing to share your thoughts 🙂</p>",
  "messages": [
    {
      "id": "3390593",
      "postDate": "01/13/2026 13:26:52",
      "content": "<p>Hi everyone 👋,</p>\n<p>I wanted to share how I’m currently constructing pixel-level ground truth masks from the provided voltage CSVs, and also hear how others are thinking about designing targets for this task.</p>\n<p><strong>Converting voltage signals into pixel-level ground truth</strong></p>\n<p>The competition provides voltage time series per lead, but the models operate on images. So the first step is converting the voltage data into pixel coordinates aligned with the rendered ECG images.</p>\n<p>What I’m doing is:</p>\n<ul>\n<li>Converting each CSV into .dat and .hea files.</li>\n<li>Using ecg_image_kit to render the ECG and extract pixel coordinates.</li>\n</ul>\n<p>For example:</p>\n<pre><code>python gen_ecg_image_from_data.py \\\n  -i root/physionet-ecg-image-digitization/train/7663343/7663343.dat \\\n  -hea root/physionet-ecg-image-digitization/train/7663343/7663343.hea \\\n  -o /root/test \\\n  -st 0 \\\n  --print_header \\\n  --lead_bbox \\\n  --lead_name_bbox \\\n  --store_config 2 \\\n  --calibration_pulse 1\n</code></pre>\n<p>This produces a JSON file with plotted_pixels per lead, which gives the exact rasterized location of the ECG traces in the image.\n(Alternatively, it’s possible to reconstruct the pixel positions directly by following the rendering logic from ecg_image_kit: lead layout + scaling of ~78.74 pixels per mV at 200 DPI.)</p>\n<p><strong>Building the segmentation mask</strong></p>\n<p>From the extracted pixel coordinates, I generate a multi-channel mask:</p>\n<ul>\n<li>Resolution: 1700 × 2200 (same as the input image)</li>\n<li>Channels: 12 leads + long lead II + background</li>\n</ul>\n<p>I draw each lead trace as a thin line using either:\n-a 1-pixel binary line, or\n-an antialiased line (e.g. matplotlib with linewidth ≈ 0.75).</p>\n<p>Optionally, I upscale the width (e.g. ×2) before training to improve temporal resolution.</p>\n<p>This gives me a clean pixel-level target aligned with the ECG image geometry.</p>\n<p><strong>Training and extracting the signal</strong></p>\n<p>I train a standard segmentation network (U-Net style), then extract the signal per column using argmax, soft-argmax, or other methods. Finally I convert pixel positions back to voltage using the known scale (≈78.74 px = 1 mV), resample to the target frequency, and evaluate SNR. This pipeline is reasonably stable and gives a consistent baseline (around ~14 dB on the leaderboard for me so far).</p>\n<p><strong>A question for discussion</strong></p>\n<p>I’ve seen people mention ideas like “high-resolution targets”, “pixel nudging”, or “refining the labels with a network”, and I’d be very interested to hear how others are approaching this part. I have done some attempts but cannot go beyond 25-30dB barrier as ground-truth segmentation. What I’m especially curious about is how others are thinking about the design of the ground truth itself.</p>\n<p>Thanks for reading — and thanks in advance if you’re willing to share your thoughts 🙂</p>",
      "rawMarkdown": "Hi everyone 👋,\n\nI wanted to share how I’m currently constructing pixel-level ground truth masks from the provided voltage CSVs, and also hear how others are thinking about designing targets for this task.\n\n**Converting voltage signals into pixel-level ground truth**\n\nThe competition provides voltage time series per lead, but the models operate on images. So the first step is converting the voltage data into pixel coordinates aligned with the rendered ECG images.\n\nWhat I’m doing is:\n\n- Converting each CSV into .dat and .hea files.\n- Using ecg_image_kit to render the ECG and extract pixel coordinates.\n\nFor example:\n```python\npython gen_ecg_image_from_data.py \\\n  -i root/physionet-ecg-image-digitization/train/7663343/7663343.dat \\\n  -hea root/physionet-ecg-image-digitization/train/7663343/7663343.hea \\\n  -o /root/test \\\n  -st 0 \\\n  --print_header \\\n  --lead_bbox \\\n  --lead_name_bbox \\\n  --store_config 2 \\\n  --calibration_pulse 1\n```\nThis produces a JSON file with plotted_pixels per lead, which gives the exact rasterized location of the ECG traces in the image.\n(Alternatively, it’s possible to reconstruct the pixel positions directly by following the rendering logic from ecg_image_kit: lead layout + scaling of ~78.74 pixels per mV at 200 DPI.)\n\n**Building the segmentation mask**\n\nFrom the extracted pixel coordinates, I generate a multi-channel mask:\n- Resolution: 1700 × 2200 (same as the input image)\n- Channels: 12 leads + long lead II + background\n\nI draw each lead trace as a thin line using either:\n-a 1-pixel binary line, or\n-an antialiased line (e.g. matplotlib with linewidth ≈ 0.75).\n\nOptionally, I upscale the width (e.g. ×2) before training to improve temporal resolution.\n\nThis gives me a clean pixel-level target aligned with the ECG image geometry.\n\n**Training and extracting the signal**\n\nI train a standard segmentation network (U-Net style), then extract the signal per column using argmax, soft-argmax, or other methods. Finally I convert pixel positions back to voltage using the known scale (≈78.74 px = 1 mV), resample to the target frequency, and evaluate SNR. This pipeline is reasonably stable and gives a consistent baseline (around ~14 dB on the leaderboard for me so far).\n\n**A question for discussion**\n\nI’ve seen people mention ideas like “high-resolution targets”, “pixel nudging”, or “refining the labels with a network”, and I’d be very interested to hear how others are approaching this part. I have done some attempts but cannot go beyond 25-30dB barrier as ground-truth segmentation. What I’m especially curious about is how others are thinking about the design of the ground truth itself.\n\nThanks for reading — and thanks in advance if you’re willing to share your thoughts 🙂",
      "votes": null
    },
    {
      "id": "3390594",
      "postDate": "01/13/2026 13:31:24",
      "content": "<p>Thank you for sharing your method. I'm stuck at this step as well. Does anyone have a better method to create a better GT mask for segmentation?</p>",
      "rawMarkdown": "Thank you for sharing your method. I'm stuck at this step as well. Does anyone have a better method to create a better GT mask for segmentation?",
      "votes": null
    },
    {
      "id": "3390606",
      "postDate": "01/13/2026 13:51:40",
      "content": "<p>forgot to mention for the above set up, I train on original image type 0001, do heavy augmentations to mimic other cases. At inference, I normalized different types to the mimic type 0001, the model prediction is not bad, but since the ground truth does not yield great score, the predicted segmented mask has even lower score… </p>",
      "rawMarkdown": "forgot to mention for the above set up, I train on original image type 0001, do heavy augmentations to mimic other cases. At inference, I normalized different types to the mimic type 0001, the model prediction is not bad, but since the ground truth does not yield great score, the predicted segmented mask has even lower score...",
      "votes": null
    },
    {
      "id": "3391702",
      "postDate": "01/15/2026 13:48:38",
      "content": "<p>i used cv2.plotlines and it worked well so far, CV do match LB. I wonder what loss function are you using?</p>",
      "rawMarkdown": "i used cv2.plotlines and it worked well so far, CV do match LB. I wonder what loss function are you using?",
      "votes": null
    },
    {
      "id": "3391919",
      "postDate": "01/15/2026 23:41:11",
      "content": "<p>Thanks for sharing. \nI am quite curious as to how you created a segmentation binary mask with antialiasing.</p>\n<p>The antialiased lines would not necessarily be a binary mask anymore would it? It would have varying pixel values. And if you applied thresholding to make it binary, it would look jagged (the issue that antialiasing attempts to fix in the first place).</p>\n<p>If the mask is indeed jagged, I am guessing this is what you mean when you say the \"ground truth does not yield great score\"</p>\n<p>In any case, I generate my own targets and training samples like I described in this notebook: <a href=\"https://www.kaggle.com/code/henrychibueze/synthesize-ecg-samples\" target=\"_blank\">https://www.kaggle.com/code/henrychibueze/synthesize-ecg-samples</a>. Training on these targets gives me a CV score of ~16.52 and LB of ~15.25. In my approach the signals are combined into a single mask, so I do use an algorithm to seperate the rows and trace them individually</p>\n<p>Of course the targets are still jagged, since they have to be binary, I have not attempted making the targets only occupy a single pixel per column.</p>",
      "rawMarkdown": "Thanks for sharing. \nI am quite curious as to how you created a segmentation binary mask with antialiasing.\n\nThe antialiased lines would not necessarily be a binary mask anymore would it? It would have varying pixel values. And if you applied thresholding to make it binary, it would look jagged (the issue that antialiasing attempts to fix in the first place).\n\nIf the mask is indeed jagged, I am guessing this is what you mean when you say the \"ground truth does not yield great score\"\n\nIn any case, I generate my own targets and training samples like I described in this notebook: https://www.kaggle.com/code/henrychibueze/synthesize-ecg-samples. Training on these targets gives me a CV score of ~16.52 and LB of ~15.25. In my approach the signals are combined into a single mask, so I do use an algorithm to seperate the rows and trace them individually\n\nOf course the targets are still jagged, since they have to be binary, I have not attempted making the targets only occupy a single pixel per column.",
      "votes": null
    },
    {
      "id": "3392108",
      "postDate": "01/16/2026 09:49:30",
      "content": "<p>Just curious, what is the SNR of your GT mask that helped you to achieve 15.25 LB? And did your training include anything special? Loss or architecture? Thank you</p>",
      "rawMarkdown": "Just curious, what is the SNR of your GT mask that helped you to achieve 15.25 LB? And did your training include anything special? Loss or architecture? Thank you",
      "votes": null
    },
    {
      "id": "3392112",
      "postDate": "01/16/2026 09:57:24",
      "content": "<p>for the segmentation model, I use normal segmentation loss (dice, bce) and eventually some regularizers for the predicted line to be center and thin but the score is not high so I don't use it for the moment</p>",
      "rawMarkdown": "for the segmentation model, I use normal segmentation loss (dice, bce) and eventually some regularizers for the predicted line to be center and thin but the score is not high so I don't use it for the moment",
      "votes": null
    },
    {
      "id": "3392123",
      "postDate": "01/16/2026 10:15:30",
      "content": "<p>Thanks! I have looked into your notebook a few times but not have yet tried it (due to other optimization part). </p>\n<p>I have tried with pillow draw 1 line thickness and LB ~ 14. For the antialiased, my idea initially is to mimic how ecg-image-kit plot in their ecg record. I have not tried training a version with the antialiased lines though.  For antialiased, i have tried to increase the resolution, so I have something like 26-27 db for the ground truth (thresholding or not), but not anymore further… and the ground truth line is so thin (high resolution) that the training will be much harder… So i am really curious how people can go up to 50 db in GT design (some discussion talking about it). and since the GT cannot have big score, I don't dive more into training these.</p>\n<p>For the mask, I trained mostly on binary ones with normal segmentation losses. Nevertheless, another idea is to train with mse loss on the varying pixel values directly (with more dilation maybe, the jagged are too much of noises), but not yet tried those. </p>",
      "rawMarkdown": "Thanks! I have looked into your notebook a few times but not have yet tried it (due to other optimization part). \n\nI have tried with pillow draw 1 line thickness and LB ~ 14. For the antialiased, my idea initially is to mimic how ecg-image-kit plot in their ecg record. I have not tried training a version with the antialiased lines though.  For antialiased, i have tried to increase the resolution, so I have something like 26-27 db for the ground truth (thresholding or not), but not anymore further... and the ground truth line is so thin (high resolution) that the training will be much harder... So i am really curious how people can go up to 50 db in GT design (some discussion talking about it). and since the GT cannot have big score, I don't dive more into training these.\n\nFor the mask, I trained mostly on binary ones with normal segmentation losses. Nevertheless, another idea is to train with mse loss on the varying pixel values directly (with more dilation maybe, the jagged are too much of noises), but not yet tried those.",
      "votes": null
    },
    {
      "id": "3392159",
      "postDate": "01/16/2026 11:16:47",
      "content": "<p>me too, my model peaked at 10.21 on CV and 13 on LB. But the public model doesn't have very good dice loss so i think they used something else</p>",
      "rawMarkdown": "me too, my model peaked at 10.21 on CV and 13 on LB. But the public model doesn't have very good dice loss so i think they used something else",
      "votes": null
    },
    {
      "id": "3392184",
      "postDate": "01/16/2026 12:10:03",
      "content": "<p>I have not bothered to check the SNR of the ground truth masks of the synthetic samples yet, only the SNR of the predictions on the competition's training images, with a model trained on the synthetic samples, which is ~16.52 db. That being said, I doubt the ground truth SNR is up to 20 db honestly, I'll have to check it later when I'm free.</p>\n<p>Another thing I'm contemplating on is the tracing algorithm, I'm currently using two algorithms, column-wise argmax (or soft-argmax) and Viterbi's algorithm (I explained how I use it <a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/662842#3378328\" target=\"_blank\">HERE</a>). The column-wise argmax  works better than the Viterbi's algorithm.</p>\n<p>I'm wondering if there's an even better way to trace the signals or if the tracer is already very good and what I actually need to work on is my ground truth masks and segmentation model. </p>\n<p>I'm also constrained to kaggle computing, it serves me well enough, but increasing image resolution would be any harder for me.</p>\n<p>I am using the Unet++ architecture and my loss functions are BCE and Tversky Loss. I trained two models, one for grid lines and grid point detection (which are used for image rectification) and the other is the signal segmentation network.</p>",
      "rawMarkdown": "I have not bothered to check the SNR of the ground truth masks of the synthetic samples yet, only the SNR of the predictions on the competition's training images, with a model trained on the synthetic samples, which is ~16.52 db. That being said, I doubt the ground truth SNR is up to 20 db honestly, I'll have to check it later when I'm free.\n\nAnother thing I'm contemplating on is the tracing algorithm, I'm currently using two algorithms, column-wise argmax (or soft-argmax) and Viterbi's algorithm (I explained how I use it [HERE](https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/662842#3378328)). The column-wise argmax  works better than the Viterbi's algorithm.\n\nI'm wondering if there's an even better way to trace the signals or if the tracer is already very good and what I actually need to work on is my ground truth masks and segmentation model. \n\nI'm also constrained to kaggle computing, it serves me well enough, but increasing image resolution would be any harder for me.\n\nI am using the Unet++ architecture and my loss functions are BCE and Tversky Loss. I trained two models, one for grid lines and grid point detection (which are used for image rectification) and the other is the signal segmentation network.",
      "votes": null
    },
    {
      "id": "3392193",
      "postDate": "01/16/2026 12:24:36",
      "content": "<p>I see.</p>\n<p>I use cv2.polylines to draw the target signals, with thickness of 1 as well, the signal is smooth until I decide to threshold (to make it binary), then it gets somewhat jagged. Then again, I thinking, perhaps I need to make a target probability field and not necessarily a target binary mask, then I can try to make the model in question approximate that probability map with a KL divergence loss. I wonder…</p>\n<p>Making the signals very thin (just one pixel) doesn't really seem like it would help, if anything it would just make the models struggle a lot.</p>\n<p>My concern with increasing image resolution is this: how would that differ from using the original resolution as is? Increasing the image resolution would be to just upsample it, as well as the binary mask with bilinear interpolation. How would this differ from training with the original resolution and then up-sampling the predicted probability scores before tracing? I wonder…</p>",
      "rawMarkdown": "I see.\n\nI use cv2.polylines to draw the target signals, with thickness of 1 as well, the signal is smooth until I decide to threshold (to make it binary), then it gets somewhat jagged. Then again, I thinking, perhaps I need to make a target probability field and not necessarily a target binary mask, then I can try to make the model in question approximate that probability map with a KL divergence loss. I wonder...\n\nMaking the signals very thin (just one pixel) doesn't really seem like it would help, if anything it would just make the models struggle a lot.\n\nMy concern with increasing image resolution is this: how would that differ from using the original resolution as is? Increasing the image resolution would be to just upsample it, as well as the binary mask with bilinear interpolation. How would this differ from training with the original resolution and then up-sampling the predicted probability scores before tracing? I wonder...",
      "votes": null
    },
    {
      "id": "3392388",
      "postDate": "01/16/2026 21:28:03",
      "content": "<p>Totally agree on making the signal thinner would make the models struggle in training :) But when testing for the ground truth (I am sure my way is not the correct way, but I try to squeeze out some SNR), the thinner the line (and we can go even thinner if we draw the line in a higher resolution image), I have better snr when extracting the series values from it. So I think I need to make the model learning to predict confidently the thin line within the center; which is kinda hard if it's too thin.. and this is where I also wonder on making a target probability field… But maybe the postprocessing can actually help to find these center lines with a sub pixel precision… so I move to that way rather. </p>\n<p>I have tried the idea of some deep net to tweak the pixel position for the GT mask a little bit so it still stay somewhat a line and somewhat coherent to the original image for better snr, but no success attempts so far. </p>\n<p>Looking forward for any suggestion and the winning solutions in a week.  </p>",
      "rawMarkdown": "Totally agree on making the signal thinner would make the models struggle in training :) But when testing for the ground truth (I am sure my way is not the correct way, but I try to squeeze out some SNR), the thinner the line (and we can go even thinner if we draw the line in a higher resolution image), I have better snr when extracting the series values from it. So I think I need to make the model learning to predict confidently the thin line within the center; which is kinda hard if it's too thin.. and this is where I also wonder on making a target probability field... But maybe the postprocessing can actually help to find these center lines with a sub pixel precision... so I move to that way rather. \n\nI have tried the idea of some deep net to tweak the pixel position for the GT mask a little bit so it still stay somewhat a line and somewhat coherent to the original image for better snr, but no success attempts so far. \n\nLooking forward for any suggestion and the winning solutions in a week.",
      "votes": null
    },
    {
      "id": "3392417",
      "postDate": "01/16/2026 23:30:48",
      "content": "<p>So, I ran some tests. <a href=\"https://www.kaggle.com/banhmimatong\" target=\"_blank\">@banhmimatong</a> </p>\n<p>With the synthesizer I made (which I have just made some local updates to), with antialiasing (no thresholding), and plotting signals with cv2.polylines (thickness = 1):</p>\n<p>at scale = 1 (resolution of 1720 x 2240), average SNR is ~17 db</p>\n<p>at scale = 2 (resolution of 3440 x 4480), average SNR is ~24 db</p>\n<p>at scale = 3 (resolution of 5160 x 6720), average SNR is ~ 30 db</p>\n<p>So it seems making high resolution plots actually creates better segments with higher SNR. And it also seems binary masks are not ideal targets, your targets would be a probability distribution, because of the antialiasing feature</p>",
      "rawMarkdown": "So, I ran some tests. @banhmimatong \n\nWith the synthesizer I made (which I have just made some local updates to), with antialiasing (no thresholding), and plotting signals with cv2.polylines (thickness = 1):\n\nat scale = 1 (resolution of 1720 x 2240), average SNR is ~17 db\n\nat scale = 2 (resolution of 3440 x 4480), average SNR is ~24 db\n\nat scale = 3 (resolution of 5160 x 6720), average SNR is ~ 30 db\n\nSo it seems making high resolution plots actually creates better segments with higher SNR. And it also seems binary masks are not ideal targets, your targets would be a probability distribution, because of the antialiasing feature",
      "votes": null
    },
    {
      "id": "3393555",
      "postDate": "01/19/2026 10:00:59",
      "content": "<p>yes that should be likely the case for me, with the public extractor argmax for postprocessing</p>",
      "rawMarkdown": "yes that should be likely the case for me, with the public extractor argmax for postprocessing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3390594,
      "author_name": "vandongtran",
      "author_url": "",
      "post_date": "01/13/2026 13:31:24",
      "content": "<p>Thank you for sharing your method. I'm stuck at this step as well. Does anyone have a better method to create a better GT mask for segmentation?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3390606,
      "author_name": "banhmimatong",
      "author_url": "",
      "post_date": "01/13/2026 13:51:40",
      "content": "<p>forgot to mention for the above set up, I train on original image type 0001, do heavy augmentations to mimic other cases. At inference, I normalized different types to the mimic type 0001, the model prediction is not bad, but since the ground truth does not yield great score, the predicted segmented mask has even lower score… </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3391702,
      "author_name": "llkh0a",
      "author_url": "",
      "post_date": "01/15/2026 13:48:38",
      "content": "<p>i used cv2.plotlines and it worked well so far, CV do match LB. I wonder what loss function are you using?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3392112,
          "author_name": "banhmimatong",
          "author_url": "",
          "post_date": "01/16/2026 09:57:24",
          "content": "<p>for the segmentation model, I use normal segmentation loss (dice, bce) and eventually some regularizers for the predicted line to be center and thin but the score is not high so I don't use it for the moment</p>",
          "votes": null,
          "replies": [
            {
              "id": 3392159,
              "author_name": "llkh0a",
              "author_url": "",
              "post_date": "01/16/2026 11:16:47",
              "content": "<p>me too, my model peaked at 10.21 on CV and 13 on LB. But the public model doesn't have very good dice loss so i think they used something else</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3391919,
      "author_name": "henrychibueze",
      "author_url": "",
      "post_date": "01/15/2026 23:41:11",
      "content": "<p>Thanks for sharing. \nI am quite curious as to how you created a segmentation binary mask with antialiasing.</p>\n<p>The antialiased lines would not necessarily be a binary mask anymore would it? It would have varying pixel values. And if you applied thresholding to make it binary, it would look jagged (the issue that antialiasing attempts to fix in the first place).</p>\n<p>If the mask is indeed jagged, I am guessing this is what you mean when you say the \"ground truth does not yield great score\"</p>\n<p>In any case, I generate my own targets and training samples like I described in this notebook: <a href=\"https://www.kaggle.com/code/henrychibueze/synthesize-ecg-samples\" target=\"_blank\">https://www.kaggle.com/code/henrychibueze/synthesize-ecg-samples</a>. Training on these targets gives me a CV score of ~16.52 and LB of ~15.25. In my approach the signals are combined into a single mask, so I do use an algorithm to seperate the rows and trace them individually</p>\n<p>Of course the targets are still jagged, since they have to be binary, I have not attempted making the targets only occupy a single pixel per column.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3392108,
          "author_name": "vandongtran",
          "author_url": "",
          "post_date": "01/16/2026 09:49:30",
          "content": "<p>Just curious, what is the SNR of your GT mask that helped you to achieve 15.25 LB? And did your training include anything special? Loss or architecture? Thank you</p>",
          "votes": null,
          "replies": [
            {
              "id": 3392184,
              "author_name": "henrychibueze",
              "author_url": "",
              "post_date": "01/16/2026 12:10:03",
              "content": "<p>I have not bothered to check the SNR of the ground truth masks of the synthetic samples yet, only the SNR of the predictions on the competition's training images, with a model trained on the synthetic samples, which is ~16.52 db. That being said, I doubt the ground truth SNR is up to 20 db honestly, I'll have to check it later when I'm free.</p>\n<p>Another thing I'm contemplating on is the tracing algorithm, I'm currently using two algorithms, column-wise argmax (or soft-argmax) and Viterbi's algorithm (I explained how I use it <a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/662842#3378328\" target=\"_blank\">HERE</a>). The column-wise argmax  works better than the Viterbi's algorithm.</p>\n<p>I'm wondering if there's an even better way to trace the signals or if the tracer is already very good and what I actually need to work on is my ground truth masks and segmentation model. </p>\n<p>I'm also constrained to kaggle computing, it serves me well enough, but increasing image resolution would be any harder for me.</p>\n<p>I am using the Unet++ architecture and my loss functions are BCE and Tversky Loss. I trained two models, one for grid lines and grid point detection (which are used for image rectification) and the other is the signal segmentation network.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3392123,
          "author_name": "banhmimatong",
          "author_url": "",
          "post_date": "01/16/2026 10:15:30",
          "content": "<p>Thanks! I have looked into your notebook a few times but not have yet tried it (due to other optimization part). </p>\n<p>I have tried with pillow draw 1 line thickness and LB ~ 14. For the antialiased, my idea initially is to mimic how ecg-image-kit plot in their ecg record. I have not tried training a version with the antialiased lines though.  For antialiased, i have tried to increase the resolution, so I have something like 26-27 db for the ground truth (thresholding or not), but not anymore further… and the ground truth line is so thin (high resolution) that the training will be much harder… So i am really curious how people can go up to 50 db in GT design (some discussion talking about it). and since the GT cannot have big score, I don't dive more into training these.</p>\n<p>For the mask, I trained mostly on binary ones with normal segmentation losses. Nevertheless, another idea is to train with mse loss on the varying pixel values directly (with more dilation maybe, the jagged are too much of noises), but not yet tried those. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3392193,
              "author_name": "henrychibueze",
              "author_url": "",
              "post_date": "01/16/2026 12:24:36",
              "content": "<p>I see.</p>\n<p>I use cv2.polylines to draw the target signals, with thickness of 1 as well, the signal is smooth until I decide to threshold (to make it binary), then it gets somewhat jagged. Then again, I thinking, perhaps I need to make a target probability field and not necessarily a target binary mask, then I can try to make the model in question approximate that probability map with a KL divergence loss. I wonder…</p>\n<p>Making the signals very thin (just one pixel) doesn't really seem like it would help, if anything it would just make the models struggle a lot.</p>\n<p>My concern with increasing image resolution is this: how would that differ from using the original resolution as is? Increasing the image resolution would be to just upsample it, as well as the binary mask with bilinear interpolation. How would this differ from training with the original resolution and then up-sampling the predicted probability scores before tracing? I wonder…</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3392388,
                  "author_name": "banhmimatong",
                  "author_url": "",
                  "post_date": "01/16/2026 21:28:03",
                  "content": "<p>Totally agree on making the signal thinner would make the models struggle in training :) But when testing for the ground truth (I am sure my way is not the correct way, but I try to squeeze out some SNR), the thinner the line (and we can go even thinner if we draw the line in a higher resolution image), I have better snr when extracting the series values from it. So I think I need to make the model learning to predict confidently the thin line within the center; which is kinda hard if it's too thin.. and this is where I also wonder on making a target probability field… But maybe the postprocessing can actually help to find these center lines with a sub pixel precision… so I move to that way rather. </p>\n<p>I have tried the idea of some deep net to tweak the pixel position for the GT mask a little bit so it still stay somewhat a line and somewhat coherent to the original image for better snr, but no success attempts so far. </p>\n<p>Looking forward for any suggestion and the winning solutions in a week.  </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3392417,
                      "author_name": "henrychibueze",
                      "author_url": "",
                      "post_date": "01/16/2026 23:30:48",
                      "content": "<p>So, I ran some tests. <a href=\"https://www.kaggle.com/banhmimatong\" target=\"_blank\">@banhmimatong</a> </p>\n<p>With the synthesizer I made (which I have just made some local updates to), with antialiasing (no thresholding), and plotting signals with cv2.polylines (thickness = 1):</p>\n<p>at scale = 1 (resolution of 1720 x 2240), average SNR is ~17 db</p>\n<p>at scale = 2 (resolution of 3440 x 4480), average SNR is ~24 db</p>\n<p>at scale = 3 (resolution of 5160 x 6720), average SNR is ~ 30 db</p>\n<p>So it seems making high resolution plots actually creates better segments with higher SNR. And it also seems binary masks are not ideal targets, your targets would be a probability distribution, because of the antialiasing feature</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3393555,
                          "author_name": "banhmimatong",
                          "author_url": "",
                          "post_date": "01/19/2026 10:00:59",
                          "content": "<p>yes that should be likely the case for me, with the public extractor argmax for postprocessing</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3390593": "Hi everyone 👋,\n\nI wanted to share how I’m currently constructing pixel-level ground truth masks from the provided voltage CSVs, and also hear how others are thinking about designing targets for this task.\n\n**Converting voltage signals into pixel-level ground truth**\n\nThe competition provides voltage time series per lead, but the models operate on images. So the first step is converting the voltage data into pixel coordinates aligned with the rendered ECG images.\n\nWhat I’m doing is:\n\n- Converting each CSV into .dat and .hea files.\n- Using ecg_image_kit to render the ECG and extract pixel coordinates.\n\nFor example:\n```python\npython gen_ecg_image_from_data.py \\\n  -i root/physionet-ecg-image-digitization/train/7663343/7663343.dat \\\n  -hea root/physionet-ecg-image-digitization/train/7663343/7663343.hea \\\n  -o /root/test \\\n  -st 0 \\\n  --print_header \\\n  --lead_bbox \\\n  --lead_name_bbox \\\n  --store_config 2 \\\n  --calibration_pulse 1\n```\nThis produces a JSON file with plotted_pixels per lead, which gives the exact rasterized location of the ECG traces in the image.\n(Alternatively, it’s possible to reconstruct the pixel positions directly by following the rendering logic from ecg_image_kit: lead layout + scaling of ~78.74 pixels per mV at 200 DPI.)\n\n**Building the segmentation mask**\n\nFrom the extracted pixel coordinates, I generate a multi-channel mask:\n- Resolution: 1700 × 2200 (same as the input image)\n- Channels: 12 leads + long lead II + background\n\nI draw each lead trace as a thin line using either:\n-a 1-pixel binary line, or\n-an antialiased line (e.g. matplotlib with linewidth ≈ 0.75).\n\nOptionally, I upscale the width (e.g. ×2) before training to improve temporal resolution.\n\nThis gives me a clean pixel-level target aligned with the ECG image geometry.\n\n**Training and extracting the signal**\n\nI train a standard segmentation network (U-Net style), then extract the signal per column using argmax, soft-argmax, or other methods. Finally I convert pixel positions back to voltage using the known scale (≈78.74 px = 1 mV), resample to the target frequency, and evaluate SNR. This pipeline is reasonably stable and gives a consistent baseline (around ~14 dB on the leaderboard for me so far).\n\n**A question for discussion**\n\nI’ve seen people mention ideas like “high-resolution targets”, “pixel nudging”, or “refining the labels with a network”, and I’d be very interested to hear how others are approaching this part. I have done some attempts but cannot go beyond 25-30dB barrier as ground-truth segmentation. What I’m especially curious about is how others are thinking about the design of the ground truth itself.\n\nThanks for reading — and thanks in advance if you’re willing to share your thoughts 🙂",
    "3390594": "Thank you for sharing your method. I'm stuck at this step as well. Does anyone have a better method to create a better GT mask for segmentation?",
    "3390606": "forgot to mention for the above set up, I train on original image type 0001, do heavy augmentations to mimic other cases. At inference, I normalized different types to the mimic type 0001, the model prediction is not bad, but since the ground truth does not yield great score, the predicted segmented mask has even lower score...",
    "3391702": "i used cv2.plotlines and it worked well so far, CV do match LB. I wonder what loss function are you using?",
    "3391919": "Thanks for sharing. \nI am quite curious as to how you created a segmentation binary mask with antialiasing.\n\nThe antialiased lines would not necessarily be a binary mask anymore would it? It would have varying pixel values. And if you applied thresholding to make it binary, it would look jagged (the issue that antialiasing attempts to fix in the first place).\n\nIf the mask is indeed jagged, I am guessing this is what you mean when you say the \"ground truth does not yield great score\"\n\nIn any case, I generate my own targets and training samples like I described in this notebook: https://www.kaggle.com/code/henrychibueze/synthesize-ecg-samples. Training on these targets gives me a CV score of ~16.52 and LB of ~15.25. In my approach the signals are combined into a single mask, so I do use an algorithm to seperate the rows and trace them individually\n\nOf course the targets are still jagged, since they have to be binary, I have not attempted making the targets only occupy a single pixel per column.",
    "3392108": "Just curious, what is the SNR of your GT mask that helped you to achieve 15.25 LB? And did your training include anything special? Loss or architecture? Thank you",
    "3392112": "for the segmentation model, I use normal segmentation loss (dice, bce) and eventually some regularizers for the predicted line to be center and thin but the score is not high so I don't use it for the moment",
    "3392123": "Thanks! I have looked into your notebook a few times but not have yet tried it (due to other optimization part). \n\nI have tried with pillow draw 1 line thickness and LB ~ 14. For the antialiased, my idea initially is to mimic how ecg-image-kit plot in their ecg record. I have not tried training a version with the antialiased lines though.  For antialiased, i have tried to increase the resolution, so I have something like 26-27 db for the ground truth (thresholding or not), but not anymore further... and the ground truth line is so thin (high resolution) that the training will be much harder... So i am really curious how people can go up to 50 db in GT design (some discussion talking about it). and since the GT cannot have big score, I don't dive more into training these.\n\nFor the mask, I trained mostly on binary ones with normal segmentation losses. Nevertheless, another idea is to train with mse loss on the varying pixel values directly (with more dilation maybe, the jagged are too much of noises), but not yet tried those.",
    "3392159": "me too, my model peaked at 10.21 on CV and 13 on LB. But the public model doesn't have very good dice loss so i think they used something else",
    "3392184": "I have not bothered to check the SNR of the ground truth masks of the synthetic samples yet, only the SNR of the predictions on the competition's training images, with a model trained on the synthetic samples, which is ~16.52 db. That being said, I doubt the ground truth SNR is up to 20 db honestly, I'll have to check it later when I'm free.\n\nAnother thing I'm contemplating on is the tracing algorithm, I'm currently using two algorithms, column-wise argmax (or soft-argmax) and Viterbi's algorithm (I explained how I use it [HERE](https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/662842#3378328)). The column-wise argmax  works better than the Viterbi's algorithm.\n\nI'm wondering if there's an even better way to trace the signals or if the tracer is already very good and what I actually need to work on is my ground truth masks and segmentation model. \n\nI'm also constrained to kaggle computing, it serves me well enough, but increasing image resolution would be any harder for me.\n\nI am using the Unet++ architecture and my loss functions are BCE and Tversky Loss. I trained two models, one for grid lines and grid point detection (which are used for image rectification) and the other is the signal segmentation network.",
    "3392193": "I see.\n\nI use cv2.polylines to draw the target signals, with thickness of 1 as well, the signal is smooth until I decide to threshold (to make it binary), then it gets somewhat jagged. Then again, I thinking, perhaps I need to make a target probability field and not necessarily a target binary mask, then I can try to make the model in question approximate that probability map with a KL divergence loss. I wonder...\n\nMaking the signals very thin (just one pixel) doesn't really seem like it would help, if anything it would just make the models struggle a lot.\n\nMy concern with increasing image resolution is this: how would that differ from using the original resolution as is? Increasing the image resolution would be to just upsample it, as well as the binary mask with bilinear interpolation. How would this differ from training with the original resolution and then up-sampling the predicted probability scores before tracing? I wonder...",
    "3392388": "Totally agree on making the signal thinner would make the models struggle in training :) But when testing for the ground truth (I am sure my way is not the correct way, but I try to squeeze out some SNR), the thinner the line (and we can go even thinner if we draw the line in a higher resolution image), I have better snr when extracting the series values from it. So I think I need to make the model learning to predict confidently the thin line within the center; which is kinda hard if it's too thin.. and this is where I also wonder on making a target probability field... But maybe the postprocessing can actually help to find these center lines with a sub pixel precision... so I move to that way rather. \n\nI have tried the idea of some deep net to tweak the pixel position for the GT mask a little bit so it still stay somewhat a line and somewhat coherent to the original image for better snr, but no success attempts so far. \n\nLooking forward for any suggestion and the winning solutions in a week.",
    "3392417": "So, I ran some tests. @banhmimatong \n\nWith the synthesizer I made (which I have just made some local updates to), with antialiasing (no thresholding), and plotting signals with cv2.polylines (thickness = 1):\n\nat scale = 1 (resolution of 1720 x 2240), average SNR is ~17 db\n\nat scale = 2 (resolution of 3440 x 4480), average SNR is ~24 db\n\nat scale = 3 (resolution of 5160 x 6720), average SNR is ~ 30 db\n\nSo it seems making high resolution plots actually creates better segments with higher SNR. And it also seems binary masks are not ideal targets, your targets would be a probability distribution, because of the antialiasing feature",
    "3393555": "yes that should be likely the case for me, with the public extractor argmax for postprocessing"
  },
  "source": "meta"
}