{
  "id": 614032,
  "title": "Segments classifier",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/614032",
  "author_name": "",
  "post_date": "2025-10-31T15:18:43.887229900Z",
  "votes": 15,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I made a public notebook <a href=\"https://www.kaggle.com/code/alejopaullier/physionet-image-multi-class-train\" target=\"_blank\">here</a> where I trained an image classifier to classify the ECG <strong>segments</strong>. Segments are variants of the same image/ECG. For example, the original ECG correspond to segment 1 and then you are provided with many additional variants, like scanned versions, stained versions, photos of the original from a monitor, etc as training data.</p>\n<p>In this competition, the training data consists of 9 different segments of each image:</p>\n<ul>\n<li><code>train/[id]/[id]-[segment].png</code>: PNG images representing the ECG signal.</li>\n<li><code>train/[id]/[id]-0001.png</code> Original color ECG image generated by ECG-image-kit.</li>\n<li><code>train/[id]/[id]-0003.png</code> Image printed in color and scanned in color.</li>\n<li><code>train/[id]/[id]-0004.png</code> Image printed in color and scanned in black and white.</li>\n<li><code>train/[id]/[id]-0005.png</code> Mobile photos of color printed images.</li>\n<li><code>train/[id]/[id]-0006.png</code> Mobile photos of ECGs on the screen of laptop.</li>\n<li><code>train/[id]/[id]-0009.png</code> Mobile photos of stained and soaked printed ECGs.</li>\n<li><code>train/[id]/[id]-0010.png</code> Mobile photos of printed ECGs with extensive damage.</li>\n<li><code>train/[id]/[id]-0011.png</code> Scans of printed ECG images with mold in color.</li>\n<li><code>train/[id]/[id]-0012.png</code> Scans of printed ECG images with mold in black and white.</li>\n</ul>\n<p>The classifier is almost perfect at predicting segments. This is expected as it is quite easy to differentiate the segments. For example, it is easy to tell which image is a photo of a computer screen (segment 6) and an original ECG (segment 1). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F1658205bbd8207e1dc64f3e8d94636bf%2Fconfusion_matrix_segments.png?generation=1761924407537758&amp;alt=media\" alt=\"\"></p>\n<p>You can use this classifier to later predict images from different kinds with different models. For instance, perhaps you have a fully dedicated model to predict ECGs from segment 6 which is super hard and the rest of segments with another model.</p>\n<p>I would say images from segments <code>1,3,4,12</code> are the easiest while the mobile photos and the stained printed ECG are the hardest. My two cents is that there are very few images from segments <code>1,3</code> in the hidden test set.</p>\n<p>Happy Kaggling :)</p>",
  "messages": [
    {
      "id": "3309380",
      "postDate": "10/31/2025 15:18:43",
      "content": "<p>I made a public notebook <a href=\"https://www.kaggle.com/code/alejopaullier/physionet-image-multi-class-train\" target=\"_blank\">here</a> where I trained an image classifier to classify the ECG <strong>segments</strong>. Segments are variants of the same image/ECG. For example, the original ECG correspond to segment 1 and then you are provided with many additional variants, like scanned versions, stained versions, photos of the original from a monitor, etc as training data.</p>\n<p>In this competition, the training data consists of 9 different segments of each image:</p>\n<ul>\n<li><code>train/[id]/[id]-[segment].png</code>: PNG images representing the ECG signal.</li>\n<li><code>train/[id]/[id]-0001.png</code> Original color ECG image generated by ECG-image-kit.</li>\n<li><code>train/[id]/[id]-0003.png</code> Image printed in color and scanned in color.</li>\n<li><code>train/[id]/[id]-0004.png</code> Image printed in color and scanned in black and white.</li>\n<li><code>train/[id]/[id]-0005.png</code> Mobile photos of color printed images.</li>\n<li><code>train/[id]/[id]-0006.png</code> Mobile photos of ECGs on the screen of laptop.</li>\n<li><code>train/[id]/[id]-0009.png</code> Mobile photos of stained and soaked printed ECGs.</li>\n<li><code>train/[id]/[id]-0010.png</code> Mobile photos of printed ECGs with extensive damage.</li>\n<li><code>train/[id]/[id]-0011.png</code> Scans of printed ECG images with mold in color.</li>\n<li><code>train/[id]/[id]-0012.png</code> Scans of printed ECG images with mold in black and white.</li>\n</ul>\n<p>The classifier is almost perfect at predicting segments. This is expected as it is quite easy to differentiate the segments. For example, it is easy to tell which image is a photo of a computer screen (segment 6) and an original ECG (segment 1). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F1658205bbd8207e1dc64f3e8d94636bf%2Fconfusion_matrix_segments.png?generation=1761924407537758&amp;alt=media\" alt=\"\"></p>\n<p>You can use this classifier to later predict images from different kinds with different models. For instance, perhaps you have a fully dedicated model to predict ECGs from segment 6 which is super hard and the rest of segments with another model.</p>\n<p>I would say images from segments <code>1,3,4,12</code> are the easiest while the mobile photos and the stained printed ECG are the hardest. My two cents is that there are very few images from segments <code>1,3</code> in the hidden test set.</p>\n<p>Happy Kaggling :)</p>",
      "rawMarkdown": "I made a public notebook [here][1] where I trained an image classifier to classify the ECG **segments**. Segments are variants of the same image/ECG. For example, the original ECG correspond to segment 1 and then you are provided with many additional variants, like scanned versions, stained versions, photos of the original from a monitor, etc as training data.\n\n In this competition, the training data consists of 9 different segments of each image:\n- `train/[id]/[id]-[segment].png`: PNG images representing the ECG signal.\n- `train/[id]/[id]-0001.png` Original color ECG image generated by ECG-image-kit.\n- `train/[id]/[id]-0003.png` Image printed in color and scanned in color.\n- `train/[id]/[id]-0004.png` Image printed in color and scanned in black and white.\n- `train/[id]/[id]-0005.png` Mobile photos of color printed images.\n- `train/[id]/[id]-0006.png` Mobile photos of ECGs on the screen of laptop.\n- `train/[id]/[id]-0009.png` Mobile photos of stained and soaked printed ECGs.\n- `train/[id]/[id]-0010.png` Mobile photos of printed ECGs with extensive damage.\n- `train/[id]/[id]-0011.png` Scans of printed ECG images with mold in color.\n- `train/[id]/[id]-0012.png` Scans of printed ECG images with mold in black and white.\n\nThe classifier is almost perfect at predicting segments. This is expected as it is quite easy to differentiate the segments. For example, it is easy to tell which image is a photo of a computer screen (segment 6) and an original ECG (segment 1). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F1658205bbd8207e1dc64f3e8d94636bf%2Fconfusion_matrix_segments.png?generation=1761924407537758&alt=media)\n\nYou can use this classifier to later predict images from different kinds with different models. For instance, perhaps you have a fully dedicated model to predict ECGs from segment 6 which is super hard and the rest of segments with another model.\n\nI would say images from segments `1,3,4,12` are the easiest while the mobile photos and the stained printed ECG are the hardest. My two cents is that there are very few images from segments `1,3` in the hidden test set.\n\nHappy Kaggling :)\n\n[1]: https://www.kaggle.com/code/alejopaullier/physionet-image-multi-class-train",
      "votes": null
    },
    {
      "id": "3309395",
      "postDate": "10/31/2025 15:52:08",
      "content": "<p>I have a model (not an image classifier) that achieves 99.3% accuracy in predicting all 9 different segments of each image, while processing over 8,700 images in just 30 seconds.</p>",
      "rawMarkdown": "I have a model (not an image classifier) that achieves 99.3% accuracy in predicting all 9 different segments of each image, while processing over 8,700 images in just 30 seconds.",
      "votes": null
    },
    {
      "id": "3309399",
      "postDate": "10/31/2025 16:19:33",
      "content": "<p>Do the 30 seconds include the time to read the 8700 images from disk?</p>",
      "rawMarkdown": "Do the 30 seconds include the time to read the 8700 images from disk?",
      "votes": null
    },
    {
      "id": "3309447",
      "postDate": "10/31/2025 17:45:35",
      "content": "<p>Nice. The task per-se is pretty easy as image segments are very distinguishable. This model takes one or two minutes I think</p>",
      "rawMarkdown": "Nice. The task per-se is pretty easy as image segments are very distinguishable. This model takes one or two minutes I think",
      "votes": null
    },
    {
      "id": "3309453",
      "postDate": "10/31/2025 18:16:50",
      "content": "<p>Yes        </p>",
      "rawMarkdown": "Yes",
      "votes": null
    },
    {
      "id": "3309457",
      "postDate": "10/31/2025 18:32:09",
      "content": "<blockquote>\n  <p>\"My two cents is that there are very few images from segments 1,3 in the hidden test set.\"</p>\n</blockquote>\n<p>Are there <em>any</em> type-1 images in the test dataset? I've generally been assuming they aren't present because any real image captured by a camera or scanner is going to be dirtier than that, so they're not suitable for calculating realistic accuracy metrics. I figured they were just provided because they could be useful to us during training, not because they'd be particularly useful during cross-validation.</p>\n<p><em>Edit: I just found a <a href=\"https://www.kaggle.com/code/guntasdhanjal/ecg-digitization-multi-resolution-lb-1-59\" target=\"_blank\">public notebook</a> where someone appears to be exploiting the presence of type-1 images, and it gave them a score improvement, so I guess the answer to my question is \"yes\", lol.</em> </p>",
      "rawMarkdown": "> \"My two cents is that there are very few images from segments 1,3 in the hidden test set.\"\n\nAre there *any* type-1 images in the test dataset? I've generally been assuming they aren't present because any real image captured by a camera or scanner is going to be dirtier than that, so they're not suitable for calculating realistic accuracy metrics. I figured they were just provided because they could be useful to us during training, not because they'd be particularly useful during cross-validation.\n\n*Edit: I just found a [public notebook](https://www.kaggle.com/code/guntasdhanjal/ecg-digitization-multi-resolution-lb-1-59) where someone appears to be exploiting the presence of type-1 images, and it gave them a score improvement, so I guess the answer to my question is \"yes\", lol.*",
      "votes": null
    },
    {
      "id": "3309464",
      "postDate": "10/31/2025 18:59:24",
      "content": "<p>I haven't probed the LB, but based on other discussions my guess is that \"easy\" images (such as types 1 and 3) are barely non-present in the test data. PhysioNet has run previous competitions in the past where these images were predicted. I think the hosts want more robust solutions in this competition, capable of predicting difficult edge cases more aligned with real world scenarios</p>",
      "rawMarkdown": "I haven't probed the LB, but based on other discussions my guess is that \"easy\" images (such as types 1 and 3) are barely non-present in the test data. PhysioNet has run previous competitions in the past where these images were predicted. I think the hosts want more robust solutions in this competition, capable of predicting difficult edge cases more aligned with real world scenarios",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3309395,
      "author_name": "omarafik",
      "author_url": "",
      "post_date": "10/31/2025 15:52:08",
      "content": "<p>I have a model (not an image classifier) that achieves 99.3% accuracy in predicting all 9 different segments of each image, while processing over 8,700 images in just 30 seconds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3309399,
          "author_name": "ambrosm",
          "author_url": "",
          "post_date": "10/31/2025 16:19:33",
          "content": "<p>Do the 30 seconds include the time to read the 8700 images from disk?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3309453,
              "author_name": "omarafik",
              "author_url": "",
              "post_date": "10/31/2025 18:16:50",
              "content": "<p>Yes        </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3309447,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "10/31/2025 17:45:35",
          "content": "<p>Nice. The task per-se is pretty easy as image segments are very distinguishable. This model takes one or two minutes I think</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3309457,
      "author_name": "jsday96",
      "author_url": "",
      "post_date": "10/31/2025 18:32:09",
      "content": "<blockquote>\n  <p>\"My two cents is that there are very few images from segments 1,3 in the hidden test set.\"</p>\n</blockquote>\n<p>Are there <em>any</em> type-1 images in the test dataset? I've generally been assuming they aren't present because any real image captured by a camera or scanner is going to be dirtier than that, so they're not suitable for calculating realistic accuracy metrics. I figured they were just provided because they could be useful to us during training, not because they'd be particularly useful during cross-validation.</p>\n<p><em>Edit: I just found a <a href=\"https://www.kaggle.com/code/guntasdhanjal/ecg-digitization-multi-resolution-lb-1-59\" target=\"_blank\">public notebook</a> where someone appears to be exploiting the presence of type-1 images, and it gave them a score improvement, so I guess the answer to my question is \"yes\", lol.</em> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3309464,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "10/31/2025 18:59:24",
          "content": "<p>I haven't probed the LB, but based on other discussions my guess is that \"easy\" images (such as types 1 and 3) are barely non-present in the test data. PhysioNet has run previous competitions in the past where these images were predicted. I think the hosts want more robust solutions in this competition, capable of predicting difficult edge cases more aligned with real world scenarios</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3309380": "I made a public notebook [here][1] where I trained an image classifier to classify the ECG **segments**. Segments are variants of the same image/ECG. For example, the original ECG correspond to segment 1 and then you are provided with many additional variants, like scanned versions, stained versions, photos of the original from a monitor, etc as training data.\n\n In this competition, the training data consists of 9 different segments of each image:\n- `train/[id]/[id]-[segment].png`: PNG images representing the ECG signal.\n- `train/[id]/[id]-0001.png` Original color ECG image generated by ECG-image-kit.\n- `train/[id]/[id]-0003.png` Image printed in color and scanned in color.\n- `train/[id]/[id]-0004.png` Image printed in color and scanned in black and white.\n- `train/[id]/[id]-0005.png` Mobile photos of color printed images.\n- `train/[id]/[id]-0006.png` Mobile photos of ECGs on the screen of laptop.\n- `train/[id]/[id]-0009.png` Mobile photos of stained and soaked printed ECGs.\n- `train/[id]/[id]-0010.png` Mobile photos of printed ECGs with extensive damage.\n- `train/[id]/[id]-0011.png` Scans of printed ECG images with mold in color.\n- `train/[id]/[id]-0012.png` Scans of printed ECG images with mold in black and white.\n\nThe classifier is almost perfect at predicting segments. This is expected as it is quite easy to differentiate the segments. For example, it is easy to tell which image is a photo of a computer screen (segment 6) and an original ECG (segment 1). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F1658205bbd8207e1dc64f3e8d94636bf%2Fconfusion_matrix_segments.png?generation=1761924407537758&alt=media)\n\nYou can use this classifier to later predict images from different kinds with different models. For instance, perhaps you have a fully dedicated model to predict ECGs from segment 6 which is super hard and the rest of segments with another model.\n\nI would say images from segments `1,3,4,12` are the easiest while the mobile photos and the stained printed ECG are the hardest. My two cents is that there are very few images from segments `1,3` in the hidden test set.\n\nHappy Kaggling :)\n\n[1]: https://www.kaggle.com/code/alejopaullier/physionet-image-multi-class-train",
    "3309395": "I have a model (not an image classifier) that achieves 99.3% accuracy in predicting all 9 different segments of each image, while processing over 8,700 images in just 30 seconds.",
    "3309399": "Do the 30 seconds include the time to read the 8700 images from disk?",
    "3309447": "Nice. The task per-se is pretty easy as image segments are very distinguishable. This model takes one or two minutes I think",
    "3309453": "Yes",
    "3309457": "> \"My two cents is that there are very few images from segments 1,3 in the hidden test set.\"\n\nAre there *any* type-1 images in the test dataset? I've generally been assuming they aren't present because any real image captured by a camera or scanner is going to be dirtier than that, so they're not suitable for calculating realistic accuracy metrics. I figured they were just provided because they could be useful to us during training, not because they'd be particularly useful during cross-validation.\n\n*Edit: I just found a [public notebook](https://www.kaggle.com/code/guntasdhanjal/ecg-digitization-multi-resolution-lb-1-59) where someone appears to be exploiting the presence of type-1 images, and it gave them a score improvement, so I guess the answer to my question is \"yes\", lol.*",
    "3309464": "I haven't probed the LB, but based on other discussions my guess is that \"easy\" images (such as types 1 and 3) are barely non-present in the test data. PhysioNet has run previous competitions in the past where these images were predicted. I think the hosts want more robust solutions in this competition, capable of predicting difficult edge cases more aligned with real world scenarios"
  },
  "source": "meta"
}