{
  "id": 670262,
  "title": "10th Place Solution - Yolo+Unet",
  "url": "/competitions/physionet-ecg-image-digitization/writeups/top-10-solution-giba",
  "author_name": "",
  "post_date": "2026-01-28T03:00:11.370Z",
  "votes": 16,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thank you to Kaggle and the sponsors for this amazing and fun competition. It was a great learning experience. I dedicated some time at the beginning and end of the competition, but unfortunately I had to stop developing models 10 days before the end to go on vacation, so I will be brief.</p>\n<p>My solution is composed by two approaches using a slightly different 3 stage pipeline each.\nBoth approaches uses a keypoints detection with a linear transformation, a grid rectification and a signal extraction stage. </p>\n<p>The main difference between the approaches is that one rectify the images based in the intersection points of the horizontal and vertical lines, plus delaunay triangularization to translate the points to the original positions. The second rectification approach uses a segmentation Unet multilabel model trained to enumerate all vertical and horizontal grid lines and given the lines segment masks finds the position of each line pixel to be translated to the original position. </p>\n<ol>\n<li><p>Keypoint detection and image transformation:\nThis stage detect 22 important keypoints in the image using yolo v12. Using the detected position of the keypoints is possible to estimate the Homography matrix to move, rotate and scale the keypoints to best match the ideal position of each keypoint. The ideal position is given by the position each keypoint is placed in images of type 0001 (original). So basically this stage automatically rotate the image and place the keypoints in the closest position as possible without distorting the image. As this is a linear transformation it will preserve all distortions in the image. This stage is done mostly to rotate the image and place all keypoints approximately in the correct position.\nTo train the yolo keypoint detector, basically, I used all 977 images of type 0001 and applied several augmentations to simulate all image types, scales and rotations as possible. This surprisingly worked very well to detect keypoints from all image types. Also the homography + warpPerspective transformation still works very well even if some keypoints weren't detected.</p></li>\n<li><p>Grid Rectification:\nThe goal of this stage is to remove all the spatial distortions from the image. To achieve this I trained a unet multilabel segmentation model to detect the masks of all the 54 vertical and 42 horizontal lines in the image. Each line was indexed with a label, so the Unet predicts the line number and the corresponding segmentation mask of each line, this makes possible to map each pixel to the original, non-distorted position. \nIn the beggining of the challenge I used a simple 2 channel Unet segmentation model to segment all vertical lines in a channel and all horizontal in the other channel. Given the prediction mask I calculated all intersection points of horizontal and vertical grid lines and enumerate according to the spatial position. This provided a grid map that I used to create triangles using delaunay algorithm, then just moved each 3 points to the correct position of type 0001 images.\nClose to the end of the challenge I invested some time to create an even better rectification algorithm and came with the grid lines multilabel approach, this improved my results in public LB from 18dB range to 21dB range </p></li>\n<li><p>Signal Extraction:\nEach ECG image contains 16 small segments, what I did is split each full image in 16 smaller images, each with a shape of 640x984 and named each image with the corresponding segment name: I, II, III, avR, avL, V1, etc. So each image gives 16 new smaller segment images that I feed to an Unet model that predicts the ECG signal mask, then a vertical argmax is applied to extract the position of the most probable vertical pixel for each x position. Using smaller images is better to manage training batch size and avoid using extremely large images.</p></li>\n</ol>\n<p>ps. since I'm still on vacation, I will update this text periodically.</p>",
  "messages": [
    {
      "id": "3397352",
      "postDate": "01/27/2026 04:23:35",
      "content": "<p>Thank you to Kaggle and the sponsors for this amazing and fun competition. It was a great learning experience. I dedicated some time at the beginning and end of the competition, but unfortunately I had to stop developing models 10 days before the end to go on vacation, so I will be brief.</p>\n<p>My solution is composed by two approaches using a slightly different 3 stage pipeline each.\nBoth approaches uses a keypoints detection with a linear transformation, a grid rectification and a signal extraction stage. </p>\n<p>The main difference between the approaches is that one rectify the images based in the intersection points of the horizontal and vertical lines, plus delaunay triangularization to translate the points to the original positions. The second rectification approach uses a segmentation Unet multilabel model trained to enumerate all vertical and horizontal grid lines and given the lines segment masks finds the position of each line pixel to be translated to the original position. </p>\n<ol>\n<li><p>Keypoint detection and image transformation:\nThis stage detect 22 important keypoints in the image using yolo v12. Using the detected position of the keypoints is possible to estimate the Homography matrix to move, rotate and scale the keypoints to best match the ideal position of each keypoint. The ideal position is given by the position each keypoint is placed in images of type 0001 (original). So basically this stage automatically rotate the image and place the keypoints in the closest position as possible without distorting the image. As this is a linear transformation it will preserve all distortions in the image. This stage is done mostly to rotate the image and place all keypoints approximately in the correct position.\nTo train the yolo keypoint detector, basically, I used all 977 images of type 0001 and applied several augmentations to simulate all image types, scales and rotations as possible. This surprisingly worked very well to detect keypoints from all image types. Also the homography + warpPerspective transformation still works very well even if some keypoints weren't detected.</p></li>\n<li><p>Grid Rectification:\nThe goal of this stage is to remove all the spatial distortions from the image. To achieve this I trained a unet multilabel segmentation model to detect the masks of all the 54 vertical and 42 horizontal lines in the image. Each line was indexed with a label, so the Unet predicts the line number and the corresponding segmentation mask of each line, this makes possible to map each pixel to the original, non-distorted position. \nIn the beggining of the challenge I used a simple 2 channel Unet segmentation model to segment all vertical lines in a channel and all horizontal in the other channel. Given the prediction mask I calculated all intersection points of horizontal and vertical grid lines and enumerate according to the spatial position. This provided a grid map that I used to create triangles using delaunay algorithm, then just moved each 3 points to the correct position of type 0001 images.\nClose to the end of the challenge I invested some time to create an even better rectification algorithm and came with the grid lines multilabel approach, this improved my results in public LB from 18dB range to 21dB range </p></li>\n<li><p>Signal Extraction:\nEach ECG image contains 16 small segments, what I did is split each full image in 16 smaller images, each with a shape of 640x984 and named each image with the corresponding segment name: I, II, III, avR, avL, V1, etc. So each image gives 16 new smaller segment images that I feed to an Unet model that predicts the ECG signal mask, then a vertical argmax is applied to extract the position of the most probable vertical pixel for each x position. Using smaller images is better to manage training batch size and avoid using extremely large images.</p></li>\n</ol>\n<p>ps. since I'm still on vacation, I will update this text periodically.</p>",
      "rawMarkdown": "Thank you to Kaggle and the sponsors for this amazing and fun competition. It was a great learning experience. I dedicated some time at the beginning and end of the competition, but unfortunately I had to stop developing models 10 days before the end to go on vacation, so I will be brief.\n\nMy solution is composed by two approaches using a slightly different 3 stage pipeline each.\nBoth approaches uses a keypoints detection with a linear transformation, a grid rectification and a signal extraction stage. \n\nThe main difference between the approaches is that one rectify the images based in the intersection points of the horizontal and vertical lines, plus delaunay triangularization to translate the points to the original positions. The second rectification approach uses a segmentation Unet multilabel model trained to enumerate all vertical and horizontal grid lines and given the lines segment masks finds the position of each line pixel to be translated to the original position. \n\n1. Keypoint detection and image transformation:\nThis stage detect 22 important keypoints in the image using yolo v12. Using the detected position of the keypoints is possible to estimate the Homography matrix to move, rotate and scale the keypoints to best match the ideal position of each keypoint. The ideal position is given by the position each keypoint is placed in images of type 0001 (original). So basically this stage automatically rotate the image and place the keypoints in the closest position as possible without distorting the image. As this is a linear transformation it will preserve all distortions in the image. This stage is done mostly to rotate the image and place all keypoints approximately in the correct position.\nTo train the yolo keypoint detector, basically, I used all 977 images of type 0001 and applied several augmentations to simulate all image types, scales and rotations as possible. This surprisingly worked very well to detect keypoints from all image types. Also the homography + warpPerspective transformation still works very well even if some keypoints weren't detected.\n\n2. Grid Rectification:\nThe goal of this stage is to remove all the spatial distortions from the image. To achieve this I trained a unet multilabel segmentation model to detect the masks of all the 54 vertical and 42 horizontal lines in the image. Each line was indexed with a label, so the Unet predicts the line number and the corresponding segmentation mask of each line, this makes possible to map each pixel to the original, non-distorted position. \nIn the beggining of the challenge I used a simple 2 channel Unet segmentation model to segment all vertical lines in a channel and all horizontal in the other channel. Given the prediction mask I calculated all intersection points of horizontal and vertical grid lines and enumerate according to the spatial position. This provided a grid map that I used to create triangles using delaunay algorithm, then just moved each 3 points to the correct position of type 0001 images.\nClose to the end of the challenge I invested some time to create an even better rectification algorithm and came with the grid lines multilabel approach, this improved my results in public LB from 18dB range to 21dB range \n\n3. Signal Extraction:\nEach ECG image contains 16 small segments, what I did is split each full image in 16 smaller images, each with a shape of 640x984 and named each image with the corresponding segment name: I, II, III, avR, avL, V1, etc. So each image gives 16 new smaller segment images that I feed to an Unet model that predicts the ECG signal mask, then a vertical argmax is applied to extract the position of the most probable vertical pixel for each x position. Using smaller images is better to manage training batch size and avoid using extremely large images.\n\n\nps. since I'm still on vacation, I will update this text periodically.",
      "votes": null
    },
    {
      "id": "3520979",
      "postDate": "09/04/2026 14:43:45",
      "content": "<p><strong>This isn't just a competition submission; it's a computational geometry masterclass! 🧠💎</strong></p>\n<p><strong>Huge congratulations on 10th place,</strong> <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>  ! Jumping from 18dB to 21dB using an exact multi-line (54V + 42H) indexing Unet for distortion rectification is pure magic.</p>\n<p>The strategy of breaking down 16 ECG lead channels into individual 640x984 patches before pixel-level signal extraction is extremely clean and scalable.</p>\n<p>**Have a great vacation, Grandmaster! ** Looking forward to seeing the code! 🙌🔥\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26124740%2F5d896e7a8c3a515438cff529344ce88c%2Fgiba1.jpeg?generation=1788533006619196&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "**This isn't just a competition submission; it's a computational geometry masterclass! 🧠💎**\n\n**Huge congratulations on 10th place,** @titericz  ! Jumping from 18dB to 21dB using an exact multi-line (54V + 42H) indexing Unet for distortion rectification is pure magic.\n\nThe strategy of breaking down 16 ECG lead channels into individual 640x984 patches before pixel-level signal extraction is extremely clean and scalable.\n\n**Have a great vacation, Grandmaster! ** Looking forward to seeing the code! 🙌🔥\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26124740%2F5d896e7a8c3a515438cff529344ce88c%2Fgiba1.jpeg?generation=1788533006619196&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3520979,
      "author_name": "datascikhan",
      "author_url": "",
      "post_date": "09/04/2026 14:43:45",
      "content": "<p><strong>This isn't just a competition submission; it's a computational geometry masterclass! 🧠💎</strong></p>\n<p><strong>Huge congratulations on 10th place,</strong> <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>  ! Jumping from 18dB to 21dB using an exact multi-line (54V + 42H) indexing Unet for distortion rectification is pure magic.</p>\n<p>The strategy of breaking down 16 ECG lead channels into individual 640x984 patches before pixel-level signal extraction is extremely clean and scalable.</p>\n<p>**Have a great vacation, Grandmaster! ** Looking forward to seeing the code! 🙌🔥\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26124740%2F5d896e7a8c3a515438cff529344ce88c%2Fgiba1.jpeg?generation=1788533006619196&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3397352": "Thank you to Kaggle and the sponsors for this amazing and fun competition. It was a great learning experience. I dedicated some time at the beginning and end of the competition, but unfortunately I had to stop developing models 10 days before the end to go on vacation, so I will be brief.\n\nMy solution is composed by two approaches using a slightly different 3 stage pipeline each.\nBoth approaches uses a keypoints detection with a linear transformation, a grid rectification and a signal extraction stage. \n\nThe main difference between the approaches is that one rectify the images based in the intersection points of the horizontal and vertical lines, plus delaunay triangularization to translate the points to the original positions. The second rectification approach uses a segmentation Unet multilabel model trained to enumerate all vertical and horizontal grid lines and given the lines segment masks finds the position of each line pixel to be translated to the original position. \n\n1. Keypoint detection and image transformation:\nThis stage detect 22 important keypoints in the image using yolo v12. Using the detected position of the keypoints is possible to estimate the Homography matrix to move, rotate and scale the keypoints to best match the ideal position of each keypoint. The ideal position is given by the position each keypoint is placed in images of type 0001 (original). So basically this stage automatically rotate the image and place the keypoints in the closest position as possible without distorting the image. As this is a linear transformation it will preserve all distortions in the image. This stage is done mostly to rotate the image and place all keypoints approximately in the correct position.\nTo train the yolo keypoint detector, basically, I used all 977 images of type 0001 and applied several augmentations to simulate all image types, scales and rotations as possible. This surprisingly worked very well to detect keypoints from all image types. Also the homography + warpPerspective transformation still works very well even if some keypoints weren't detected.\n\n2. Grid Rectification:\nThe goal of this stage is to remove all the spatial distortions from the image. To achieve this I trained a unet multilabel segmentation model to detect the masks of all the 54 vertical and 42 horizontal lines in the image. Each line was indexed with a label, so the Unet predicts the line number and the corresponding segmentation mask of each line, this makes possible to map each pixel to the original, non-distorted position. \nIn the beggining of the challenge I used a simple 2 channel Unet segmentation model to segment all vertical lines in a channel and all horizontal in the other channel. Given the prediction mask I calculated all intersection points of horizontal and vertical grid lines and enumerate according to the spatial position. This provided a grid map that I used to create triangles using delaunay algorithm, then just moved each 3 points to the correct position of type 0001 images.\nClose to the end of the challenge I invested some time to create an even better rectification algorithm and came with the grid lines multilabel approach, this improved my results in public LB from 18dB range to 21dB range \n\n3. Signal Extraction:\nEach ECG image contains 16 small segments, what I did is split each full image in 16 smaller images, each with a shape of 640x984 and named each image with the corresponding segment name: I, II, III, avR, avL, V1, etc. So each image gives 16 new smaller segment images that I feed to an Unet model that predicts the ECG signal mask, then a vertical argmax is applied to extract the position of the most probable vertical pixel for each x position. Using smaller images is better to manage training batch size and avoid using extremely large images.\n\n\nps. since I'm still on vacation, I will update this text periodically.",
    "3520979": "**This isn't just a competition submission; it's a computational geometry masterclass! 🧠💎**\n\n**Huge congratulations on 10th place,** @titericz  ! Jumping from 18dB to 21dB using an exact multi-line (54V + 42H) indexing Unet for distortion rectification is pure magic.\n\nThe strategy of breaking down 16 ECG lead channels into individual 640x984 patches before pixel-level signal extraction is extremely clean and scalable.\n\n**Have a great vacation, Grandmaster! ** Looking forward to seeing the code! 🙌🔥\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26124740%2F5d896e7a8c3a515438cff529344ce88c%2Fgiba1.jpeg?generation=1788533006619196&alt=media)"
  },
  "source": "meta"
}