{
  "id": 612769,
  "title": "Another data generation competition?",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/612769",
  "author_name": "",
  "post_date": "2025-10-22T01:15:01.961508300Z",
  "votes": 28,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Since data synthetic tool is provided, I guess the winner of this competition is someone who trains with more and more data. </p>\n<p>I bet it requires endless amount of computing resource to win.</p>\n<p><a href=\"https://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator\" target=\"_blank\">https://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator</a></p>",
  "messages": [
    {
      "id": "3305082",
      "postDate": "10/22/2025 01:15:01",
      "content": "<p>Since data synthetic tool is provided, I guess the winner of this competition is someone who trains with more and more data. </p>\n<p>I bet it requires endless amount of computing resource to win.</p>\n<p><a href=\"https://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator\" target=\"_blank\">https://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator</a></p>",
      "rawMarkdown": "Since data synthetic tool is provided, I guess the winner of this competition is someone who trains with more and more data. \n\nI bet it requires endless amount of computing resource to win.\n\nhttps://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator",
      "votes": null
    },
    {
      "id": "3305086",
      "postDate": "10/22/2025 01:28:05",
      "content": "<p>I feel like kaggle competition becomes a resource game.</p>",
      "rawMarkdown": "I feel like kaggle competition becomes a resource game.",
      "votes": null
    },
    {
      "id": "3305092",
      "postDate": "10/22/2025 02:01:00",
      "content": "<p>We sympathize with that sentiment and have no wish to exacerbate the rapidly increasing carbon footprint of AI. We provided <a href=\"https://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator\" target=\"_blank\">the ECG image generator you linked to</a> for our <a href=\"http://physionetchallenge.org/2024\" target=\"_blank\">previous challenge</a> in case people wanted to use it to generate artificial training data. However, the challenge doesn't allow endless compute, and therefore leveraging pretrained models, and wise choices of training data may be the key to success. There are lots of traditional techniques out there, and ways to leverage the correlated structure of the data.</p>",
      "rawMarkdown": "We sympathize with that sentiment and have no wish to exacerbate the rapidly increasing carbon footprint of AI. We provided [the ECG image generator you linked to](https://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator) for our [previous challenge](http://physionetchallenge.org/2024) in case people wanted to use it to generate artificial training data. However, the challenge doesn't allow endless compute, and therefore leveraging pretrained models, and wise choices of training data may be the key to success. There are lots of traditional techniques out there, and ways to leverage the correlated structure of the data.",
      "votes": null
    },
    {
      "id": "3305099",
      "postDate": "10/22/2025 02:51:49",
      "content": "<p>I know some competitions where we can generate as many data, and no exception, top ranker uses fudge amount to train data (~TB). More train data to win, I believe this is a gold rule of training machine leaning models.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2025\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2025</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025\" target=\"_blank\">https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025</a></li>\n</ul>\n<p>In my personal view, data generation itself is not problem, but quality of output is problem.<br>\nSuppose under certain task setup, the winner models just ended up with overfitting to competition data setup, and not useful to general data setup. This type of competition would be just a resource game, and the RoI of competition (as a whole community) is poor.</p>\n<p>I hope this competition is not this . And if it is or not, I forecast this competition's resource requirement surges if we want to win.</p>",
      "rawMarkdown": "I know some competitions where we can generate as many data, and no exception, top ranker uses fudge amount to train data (~TB). More train data to win, I believe this is a gold rule of training machine leaning models.\n\n- https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim\n- https://www.kaggle.com/competitions/ariel-data-challenge-2025\n- https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025\n\nIn my personal view, data generation itself is not problem, but quality of output is problem.\nSuppose under certain task setup, the winner models just ended up with overfitting to competition data setup, and not useful to general data setup. This type of competition would be just a resource game, and the RoI of competition (as a whole community) is poor.\n\nI hope this competition is not this . And if it is or not, I forecast this competition's resource requirement surges if we want to win.",
      "votes": null
    },
    {
      "id": "3305107",
      "postDate": "10/22/2025 03:30:07",
      "content": "<p>Am I correct to assume you used real ECG images for the test set, not ones generated by ecg-image-kit?</p>\n<p>If so, there could be an interesting distribution shift between the generated training images and the real test images. If that's the case, then creative data generation tricks might yield more score improvement than naively scaling up the training dataset.</p>",
      "rawMarkdown": "Am I correct to assume you used real ECG images for the test set, not ones generated by ecg-image-kit?\n\nIf so, there could be an interesting distribution shift between the generated training images and the real test images. If that's the case, then creative data generation tricks might yield more score improvement than naively scaling up the training dataset.",
      "votes": null
    },
    {
      "id": "3305118",
      "postDate": "10/22/2025 03:57:00",
      "content": "<p>This is a general paradigm problem with machine learning. If your test data looks very much like your training data, then you won't be penalized for failing to generalize beyond the examples, especially when there aren't millions of them. (Medical data is time consuming to generate accurate labels for, so good data sets tend to be relatively small, and large data sets tend to be quite noisy.) We have been running challenges like this for 26 years - see <a href=\"https://PhysioNetChallenges.org\" target=\"_blank\">PhysioNetChallenges.org</a> - in which we specifically address issues like this, ensuring our test data has some similar data to the training set, and some different data. In this way we can identify those entries that are able to generalize well, and those that are just lucky. In our regular challenges we also require teams to write scientific articles describing their work, and defend the work at a public forum. In addition, we require that the teams' algorithms are fully retrainable, and we retrain them with a strict limit on compute time, with tests to ensure that meaningful training actually occurs. This makes it harder to brute force a win. However, this takes enormous effort on the part of both the competitors and the organizers. The Kaggle platform gives us a unique opportunity to reach many more people, and although there's a risk that the winning team ends up not generalizing very well, we think we've built enough safeguards into this challenge to minimize that risk. </p>",
      "rawMarkdown": "This is a general paradigm problem with machine learning. If your test data looks very much like your training data, then you won't be penalized for failing to generalize beyond the examples, especially when there aren't millions of them. (Medical data is time consuming to generate accurate labels for, so good data sets tend to be relatively small, and large data sets tend to be quite noisy.) We have been running challenges like this for 26 years - see [PhysioNetChallenges.org](https://PhysioNetChallenges.org) - in which we specifically address issues like this, ensuring our test data has some similar data to the training set, and some different data. In this way we can identify those entries that are able to generalize well, and those that are just lucky. In our regular challenges we also require teams to write scientific articles describing their work, and defend the work at a public forum. In addition, we require that the teams' algorithms are fully retrainable, and we retrain them with a strict limit on compute time, with tests to ensure that meaningful training actually occurs. This makes it harder to brute force a win. However, this takes enormous effort on the part of both the competitors and the organizers. The Kaggle platform gives us a unique opportunity to reach many more people, and although there's a risk that the winning team ends up not generalizing very well, we think we've built enough safeguards into this challenge to minimize that risk.",
      "votes": null
    },
    {
      "id": "3305123",
      "postDate": "10/22/2025 04:11:07",
      "content": "<p><a href=\"https://www.kaggle.com/gdclifford\" target=\"_blank\">@gdclifford</a> </p>\n<blockquote>\n  <p>In addition, we require that the teams' algorithms are fully retrainable, and we retrain them with a strict limit on compute time</p>\n</blockquote>\n<p>Just a confirmation, does this Kaggle competition has specific rule to restrict training time?</p>",
      "rawMarkdown": "gdclifford \n\n> In addition, we require that the teams' algorithms are fully retrainable, and we retrain them with a strict limit on compute time\n\nJust a confirmation, does this Kaggle competition has specific rule to restrict training time?",
      "votes": null
    },
    {
      "id": "3305213",
      "postDate": "10/22/2025 09:09:26",
      "content": "<p>Training on massive data corpora has been an issue since around 2023, following a series of LLM competitions. But that’s fine — there are always some competitions that don’t require extensive GPU resources</p>\n<p>Recently, there was a Yale competition with a similar 'problem' (with the only diff that the tool wasn't provided, but it was relatively easy to reverse-engineer it) where teams with modest compute budgets (under $500) still managed to get gold. Of course, having GPUs is an advantage, but really smart and creative people can still win with ease</p>",
      "rawMarkdown": "Training on massive data corpora has been an issue since around 2023, following a series of LLM competitions. But that’s fine — there are always some competitions that don’t require extensive GPU resources\n\nRecently, there was a Yale competition with a similar 'problem' (with the only diff that the tool wasn't provided, but it was relatively easy to reverse-engineer it) where teams with modest compute budgets (under $500) still managed to get gold. Of course, having GPUs is an advantage, but really smart and creative people can still win with ease",
      "votes": null
    },
    {
      "id": "3305269",
      "postDate": "10/22/2025 12:10:20",
      "content": "<p>No - I'm referring to our other challenges where we have restricted both training and inference time. Kaggle competitions only restrict inference time per: <a href=\"https://www.kaggle.com/docs/notebooks\" target=\"_blank\">https://www.kaggle.com/docs/notebooks</a> (look up \"Technical Specifications\")</p>",
      "rawMarkdown": "No - I'm referring to our other challenges where we have restricted both training and inference time. Kaggle competitions only restrict inference time per: [https://www.kaggle.com/docs/notebooks](https://www.kaggle.com/docs/notebooks) (look up \"Technical Specifications\")",
      "votes": null
    },
    {
      "id": "3305275",
      "postDate": "10/22/2025 12:20:54",
      "content": "<p>10000TB data incoming</p>",
      "rawMarkdown": "10000TB data incoming",
      "votes": null
    },
    {
      "id": "3305282",
      "postDate": "10/22/2025 12:34:20",
      "content": "<p>ecg-image-kit is not a synthetic ECG generator. It prints real input ECG time-series on realistic ECG grids, just like what a standard ECG machine does. The synthetic features of ecg-image-kit are the imaging artifacts and variations like rotation and wrinkles, not the ECG itself. The shared data have all been printed out in hardcopy and distorted with real world imaging artifacts. You can read more about the dataset in the references we've shared under the challenge description.</p>",
      "rawMarkdown": "ecg-image-kit is not a synthetic ECG generator. It prints real input ECG time-series on realistic ECG grids, just like what a standard ECG machine does. The synthetic features of ecg-image-kit are the imaging artifacts and variations like rotation and wrinkles, not the ECG itself. The shared data have all been printed out in hardcopy and distorted with real world imaging artifacts. You can read more about the dataset in the references we've shared under the challenge description.",
      "votes": null
    },
    {
      "id": "3305285",
      "postDate": "10/22/2025 12:38:23",
      "content": "<p>We did not generate artificial artifacts in the training or test images using ecg-image-kit. The images in the training and test data are created in the same way - we physically print real ECGs to ECG graph paper, then scan them back in (or photograph them) per the competition description. We also add real physical artifacts (creases, writing, stains, mold, etc.) and scan them. We provide ecg-image-kit to give you a way to simulate these artifacts, but we don't guarantee that it is exactly the same as the real ECG scans. It's up to you to judge how useful this is (perhaps by cross validating on the training data), and how much external (public?) data you need to use to generate enough useful training. You can also generate some images manually, like we did, if you do not think the artifacts generated are useful enough, or use some interesting training approaches that don't overfit on the artificial data you generate. </p>",
      "rawMarkdown": "We did not generate artificial artifacts in the training or test images using ecg-image-kit. The images in the training and test data are created in the same way - we physically print real ECGs to ECG graph paper, then scan them back in (or photograph them) per the competition description. We also add real physical artifacts (creases, writing, stains, mold, etc.) and scan them. We provide ecg-image-kit to give you a way to simulate these artifacts, but we don't guarantee that it is exactly the same as the real ECG scans. It's up to you to judge how useful this is (perhaps by cross validating on the training data), and how much external (public?) data you need to use to generate enough useful training. You can also generate some images manually, like we did, if you do not think the artifacts generated are useful enough, or use some interesting training approaches that don't overfit on the artificial data you generate.",
      "votes": null
    },
    {
      "id": "3305306",
      "postDate": "10/22/2025 13:30:30",
      "content": "<blockquote>\n  <p>The images in the training and test data are created in the same way - we physically print real ECGs to ECG graph paper, then scan them back in (or photograph them)</p>\n</blockquote>\n<p>Ah, thanks for clarifying. This answers my question.</p>\n<p>I thought all the training images (including the augmented versions) came from ecg-image-kit because the data documentation says \"<em>train/[id]/[id]-0001.png Original color ECG image generated by ECG-image-kit</em>.\" However, with the additional information from your response, my current understanding is that you <em>only</em> used ecg-image-kit to generate the <em>clean</em> images; the augmented versions were created by physically printing and taking pictures of the clean images instead of using ecg-image-kit's built-in data augmentation functionality to simulate factors like wrinkles in the paper or rotation.</p>",
      "rawMarkdown": "> The images in the training and test data are created in the same way - we physically print real ECGs to ECG graph paper, then scan them back in (or photograph them)\n\nAh, thanks for clarifying. This answers my question.\n\nI thought all the training images (including the augmented versions) came from ecg-image-kit because the data documentation says \"*train/[id]/[id]-0001.png Original color ECG image generated by ECG-image-kit*.\" However, with the additional information from your response, my current understanding is that you *only* used ecg-image-kit to generate the *clean* images; the augmented versions were created by physically printing and taking pictures of the clean images instead of using ecg-image-kit's built-in data augmentation functionality to simulate factors like wrinkles in the paper or rotation.",
      "votes": null
    },
    {
      "id": "3305317",
      "postDate": "10/22/2025 13:48:27",
      "content": "<p>That is correct.</p>",
      "rawMarkdown": "That is correct.",
      "votes": null
    },
    {
      "id": "3305326",
      "postDate": "10/22/2025 14:05:59",
      "content": "<p>That's good new for me. I bought an 8TB SSD this year for kaggle; I only need to buy 1249 more!</p>",
      "rawMarkdown": "That's good new for me. I bought an 8TB SSD this year for kaggle; I only need to buy 1249 more!",
      "votes": null
    },
    {
      "id": "3305328",
      "postDate": "10/22/2025 14:13:03",
      "content": "<p>should sponsor me some as well</p>",
      "rawMarkdown": "should sponsor me some as well",
      "votes": null
    },
    {
      "id": "3305362",
      "postDate": "10/22/2025 15:17:15",
      "content": "<p>I also bought 8TB SSD. 8TB SDD has 4800TBW, so you only need 16TB SSD in the best case.</p>",
      "rawMarkdown": "I also bought 8TB SSD. 8TB SDD has 4800TBW, so you only need 16TB SSD in the best case.",
      "votes": null
    },
    {
      "id": "3305548",
      "postDate": "10/22/2025 23:43:35",
      "content": "<p>Hello,<br>\nI've recently completed my master's in data science and learnt aout AI/Ml.Also,I'm curious to know about the competition becuase it is my first time to participate can somebody help me with the rules and process where I can get an idea to proceed further.Thanks</p>",
      "rawMarkdown": "Hello,\nI've recently completed my master's in data science and learnt aout AI/Ml.Also,I'm curious to know about the competition becuase it is my first time to participate can somebody help me with the rules and process where I can get an idea to proceed further.Thanks",
      "votes": null
    },
    {
      "id": "3305679",
      "postDate": "10/23/2025 08:28:21",
      "content": "<p>And take 10000 TB of photos of moldy papers, it sure sounds realistic.</p>",
      "rawMarkdown": "And take 10000 TB of photos of moldy papers, it sure sounds realistic.",
      "votes": null
    },
    {
      "id": "3305802",
      "postDate": "10/23/2025 13:26:21",
      "content": "<p>It depends on your model architecture. Do not underestimate using, or combining your methods with, more classical signal processing and computer vision techniques. Reading and incorporating contextual knowledge regarding the ECG also goes a long way.</p>",
      "rawMarkdown": "It depends on your model architecture. Do not underestimate using, or combining your methods with, more classical signal processing and computer vision techniques. Reading and incorporating contextual knowledge regarding the ECG also goes a long way.",
      "votes": null
    },
    {
      "id": "3306145",
      "postDate": "10/24/2025 02:33:44",
      "content": "<p>Might can make a pure keypoint+alignment based method. I found that basic Orb approach can generate a lot of good points on the ecg image, then align each point with the actual timestamp in csv file would be nice. But if it's just for winning, the above 10000TB guy is still in favor of this competition.</p>",
      "rawMarkdown": "Might can make a pure keypoint+alignment based method. I found that basic Orb approach can generate a lot of good points on the ecg image, then align each point with the actual timestamp in csv file would be nice. But if it's just for winning, the above 10000TB guy is still in favor of this competition.",
      "votes": null
    },
    {
      "id": "3307092",
      "postDate": "10/26/2025 05:41:47",
      "content": "<p>But honestly, I’m curious to see if someone manages to pull a gold with clever signal processing instead of 10,000TB of synthetic ECGs.</p>",
      "rawMarkdown": "But honestly, I’m curious to see if someone manages to pull a gold with clever signal processing instead of 10,000TB of synthetic ECGs.",
      "votes": null
    },
    {
      "id": "3307106",
      "postDate": "10/26/2025 06:27:38",
      "content": "<p>People are begining to forget that $500 isn't that cheap…</p>",
      "rawMarkdown": "People are begining to forget that $500 isn't that cheap...",
      "votes": null
    },
    {
      "id": "3308138",
      "postDate": "10/28/2025 17:22:26",
      "content": "<p>true that just found this out the hard way lol</p>",
      "rawMarkdown": "true that just found this out the hard way lol",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3305086,
      "author_name": "tom99763",
      "author_url": "",
      "post_date": "10/22/2025 01:28:05",
      "content": "<p>I feel like kaggle competition becomes a resource game.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3305802,
          "author_name": "r2241272",
          "author_url": "",
          "post_date": "10/23/2025 13:26:21",
          "content": "<p>It depends on your model architecture. Do not underestimate using, or combining your methods with, more classical signal processing and computer vision techniques. Reading and incorporating contextual knowledge regarding the ECG also goes a long way.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3306145,
              "author_name": "tom99763",
              "author_url": "",
              "post_date": "10/24/2025 02:33:44",
              "content": "<p>Might can make a pure keypoint+alignment based method. I found that basic Orb approach can generate a lot of good points on the ecg image, then align each point with the actual timestamp in csv file would be nice. But if it's just for winning, the above 10000TB guy is still in favor of this competition.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3305092,
      "author_name": "gdclifford",
      "author_url": "",
      "post_date": "10/22/2025 02:01:00",
      "content": "<p>We sympathize with that sentiment and have no wish to exacerbate the rapidly increasing carbon footprint of AI. We provided <a href=\"https://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator\" target=\"_blank\">the ECG image generator you linked to</a> for our <a href=\"http://physionetchallenge.org/2024\" target=\"_blank\">previous challenge</a> in case people wanted to use it to generate artificial training data. However, the challenge doesn't allow endless compute, and therefore leveraging pretrained models, and wise choices of training data may be the key to success. There are lots of traditional techniques out there, and ways to leverage the correlated structure of the data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3305099,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "10/22/2025 02:51:49",
          "content": "<p>I know some competitions where we can generate as many data, and no exception, top ranker uses fudge amount to train data (~TB). More train data to win, I believe this is a gold rule of training machine leaning models.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2025\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2025</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025\" target=\"_blank\">https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025</a></li>\n</ul>\n<p>In my personal view, data generation itself is not problem, but quality of output is problem.<br>\nSuppose under certain task setup, the winner models just ended up with overfitting to competition data setup, and not useful to general data setup. This type of competition would be just a resource game, and the RoI of competition (as a whole community) is poor.</p>\n<p>I hope this competition is not this . And if it is or not, I forecast this competition's resource requirement surges if we want to win.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3305118,
              "author_name": "gdclifford",
              "author_url": "",
              "post_date": "10/22/2025 03:57:00",
              "content": "<p>This is a general paradigm problem with machine learning. If your test data looks very much like your training data, then you won't be penalized for failing to generalize beyond the examples, especially when there aren't millions of them. (Medical data is time consuming to generate accurate labels for, so good data sets tend to be relatively small, and large data sets tend to be quite noisy.) We have been running challenges like this for 26 years - see <a href=\"https://PhysioNetChallenges.org\" target=\"_blank\">PhysioNetChallenges.org</a> - in which we specifically address issues like this, ensuring our test data has some similar data to the training set, and some different data. In this way we can identify those entries that are able to generalize well, and those that are just lucky. In our regular challenges we also require teams to write scientific articles describing their work, and defend the work at a public forum. In addition, we require that the teams' algorithms are fully retrainable, and we retrain them with a strict limit on compute time, with tests to ensure that meaningful training actually occurs. This makes it harder to brute force a win. However, this takes enormous effort on the part of both the competitors and the organizers. The Kaggle platform gives us a unique opportunity to reach many more people, and although there's a risk that the winning team ends up not generalizing very well, we think we've built enough safeguards into this challenge to minimize that risk. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3305123,
                  "author_name": "tatamikenn",
                  "author_url": "",
                  "post_date": "10/22/2025 04:11:07",
                  "content": "<p><a href=\"https://www.kaggle.com/gdclifford\" target=\"_blank\">@gdclifford</a> </p>\n<blockquote>\n  <p>In addition, we require that the teams' algorithms are fully retrainable, and we retrain them with a strict limit on compute time</p>\n</blockquote>\n<p>Just a confirmation, does this Kaggle competition has specific rule to restrict training time?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3305269,
                      "author_name": "gdclifford",
                      "author_url": "",
                      "post_date": "10/22/2025 12:10:20",
                      "content": "<p>No - I'm referring to our other challenges where we have restricted both training and inference time. Kaggle competitions only restrict inference time per: <a href=\"https://www.kaggle.com/docs/notebooks\" target=\"_blank\">https://www.kaggle.com/docs/notebooks</a> (look up \"Technical Specifications\")</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3305107,
          "author_name": "jsday96",
          "author_url": "",
          "post_date": "10/22/2025 03:30:07",
          "content": "<p>Am I correct to assume you used real ECG images for the test set, not ones generated by ecg-image-kit?</p>\n<p>If so, there could be an interesting distribution shift between the generated training images and the real test images. If that's the case, then creative data generation tricks might yield more score improvement than naively scaling up the training dataset.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3305282,
              "author_name": "r2241272",
              "author_url": "",
              "post_date": "10/22/2025 12:34:20",
              "content": "<p>ecg-image-kit is not a synthetic ECG generator. It prints real input ECG time-series on realistic ECG grids, just like what a standard ECG machine does. The synthetic features of ecg-image-kit are the imaging artifacts and variations like rotation and wrinkles, not the ECG itself. The shared data have all been printed out in hardcopy and distorted with real world imaging artifacts. You can read more about the dataset in the references we've shared under the challenge description.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3305285,
              "author_name": "gdclifford",
              "author_url": "",
              "post_date": "10/22/2025 12:38:23",
              "content": "<p>We did not generate artificial artifacts in the training or test images using ecg-image-kit. The images in the training and test data are created in the same way - we physically print real ECGs to ECG graph paper, then scan them back in (or photograph them) per the competition description. We also add real physical artifacts (creases, writing, stains, mold, etc.) and scan them. We provide ecg-image-kit to give you a way to simulate these artifacts, but we don't guarantee that it is exactly the same as the real ECG scans. It's up to you to judge how useful this is (perhaps by cross validating on the training data), and how much external (public?) data you need to use to generate enough useful training. You can also generate some images manually, like we did, if you do not think the artifacts generated are useful enough, or use some interesting training approaches that don't overfit on the artificial data you generate. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3305306,
                  "author_name": "jsday96",
                  "author_url": "",
                  "post_date": "10/22/2025 13:30:30",
                  "content": "<blockquote>\n  <p>The images in the training and test data are created in the same way - we physically print real ECGs to ECG graph paper, then scan them back in (or photograph them)</p>\n</blockquote>\n<p>Ah, thanks for clarifying. This answers my question.</p>\n<p>I thought all the training images (including the augmented versions) came from ecg-image-kit because the data documentation says \"<em>train/[id]/[id]-0001.png Original color ECG image generated by ECG-image-kit</em>.\" However, with the additional information from your response, my current understanding is that you <em>only</em> used ecg-image-kit to generate the <em>clean</em> images; the augmented versions were created by physically printing and taking pictures of the clean images instead of using ecg-image-kit's built-in data augmentation functionality to simulate factors like wrinkles in the paper or rotation.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3305317,
                      "author_name": "r2241272",
                      "author_url": "",
                      "post_date": "10/22/2025 13:48:27",
                      "content": "<p>That is correct.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3305213,
      "author_name": "samson8",
      "author_url": "",
      "post_date": "10/22/2025 09:09:26",
      "content": "<p>Training on massive data corpora has been an issue since around 2023, following a series of LLM competitions. But that’s fine — there are always some competitions that don’t require extensive GPU resources</p>\n<p>Recently, there was a Yale competition with a similar 'problem' (with the only diff that the tool wasn't provided, but it was relatively easy to reverse-engineer it) where teams with modest compute budgets (under $500) still managed to get gold. Of course, having GPUs is an advantage, but really smart and creative people can still win with ease</p>",
      "votes": null,
      "replies": [
        {
          "id": 3307106,
          "author_name": "cnumber",
          "author_url": "",
          "post_date": "10/26/2025 06:27:38",
          "content": "<p>People are begining to forget that $500 isn't that cheap…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3305275,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "10/22/2025 12:20:54",
      "content": "<p>10000TB data incoming</p>",
      "votes": null,
      "replies": [
        {
          "id": 3305326,
          "author_name": "cnumber",
          "author_url": "",
          "post_date": "10/22/2025 14:05:59",
          "content": "<p>That's good new for me. I bought an 8TB SSD this year for kaggle; I only need to buy 1249 more!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3305328,
              "author_name": "tom99763",
              "author_url": "",
              "post_date": "10/22/2025 14:13:03",
              "content": "<p>should sponsor me some as well</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3305362,
              "author_name": "yuanzhezhou",
              "author_url": "",
              "post_date": "10/22/2025 15:17:15",
              "content": "<p>I also bought 8TB SSD. 8TB SDD has 4800TBW, so you only need 16TB SSD in the best case.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3305679,
              "author_name": "cnumber",
              "author_url": "",
              "post_date": "10/23/2025 08:28:21",
              "content": "<p>And take 10000 TB of photos of moldy papers, it sure sounds realistic.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3307092,
          "author_name": "ladiposamson",
          "author_url": "",
          "post_date": "10/26/2025 05:41:47",
          "content": "<p>But honestly, I’m curious to see if someone manages to pull a gold with clever signal processing instead of 10,000TB of synthetic ECGs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3305548,
      "author_name": "poojashinde9",
      "author_url": "",
      "post_date": "10/22/2025 23:43:35",
      "content": "<p>Hello,<br>\nI've recently completed my master's in data science and learnt aout AI/Ml.Also,I'm curious to know about the competition becuase it is my first time to participate can somebody help me with the rules and process where I can get an idea to proceed further.Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3308138,
      "author_name": "simongunyali",
      "author_url": "",
      "post_date": "10/28/2025 17:22:26",
      "content": "<p>true that just found this out the hard way lol</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3305082": "Since data synthetic tool is provided, I guess the winner of this competition is someone who trains with more and more data. \n\nI bet it requires endless amount of computing resource to win.\n\nhttps://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator",
    "3305086": "I feel like kaggle competition becomes a resource game.",
    "3305092": "We sympathize with that sentiment and have no wish to exacerbate the rapidly increasing carbon footprint of AI. We provided [the ECG image generator you linked to](https://github.com/alphanumericslab/ecg-image-kit/tree/main/codes/ecg-image-generator) for our [previous challenge](http://physionetchallenge.org/2024) in case people wanted to use it to generate artificial training data. However, the challenge doesn't allow endless compute, and therefore leveraging pretrained models, and wise choices of training data may be the key to success. There are lots of traditional techniques out there, and ways to leverage the correlated structure of the data.",
    "3305099": "I know some competitions where we can generate as many data, and no exception, top ranker uses fudge amount to train data (~TB). More train data to win, I believe this is a gold rule of training machine leaning models.\n\n- https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim\n- https://www.kaggle.com/competitions/ariel-data-challenge-2025\n- https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025\n\nIn my personal view, data generation itself is not problem, but quality of output is problem.\nSuppose under certain task setup, the winner models just ended up with overfitting to competition data setup, and not useful to general data setup. This type of competition would be just a resource game, and the RoI of competition (as a whole community) is poor.\n\nI hope this competition is not this . And if it is or not, I forecast this competition's resource requirement surges if we want to win.",
    "3305107": "Am I correct to assume you used real ECG images for the test set, not ones generated by ecg-image-kit?\n\nIf so, there could be an interesting distribution shift between the generated training images and the real test images. If that's the case, then creative data generation tricks might yield more score improvement than naively scaling up the training dataset.",
    "3305118": "This is a general paradigm problem with machine learning. If your test data looks very much like your training data, then you won't be penalized for failing to generalize beyond the examples, especially when there aren't millions of them. (Medical data is time consuming to generate accurate labels for, so good data sets tend to be relatively small, and large data sets tend to be quite noisy.) We have been running challenges like this for 26 years - see [PhysioNetChallenges.org](https://PhysioNetChallenges.org) - in which we specifically address issues like this, ensuring our test data has some similar data to the training set, and some different data. In this way we can identify those entries that are able to generalize well, and those that are just lucky. In our regular challenges we also require teams to write scientific articles describing their work, and defend the work at a public forum. In addition, we require that the teams' algorithms are fully retrainable, and we retrain them with a strict limit on compute time, with tests to ensure that meaningful training actually occurs. This makes it harder to brute force a win. However, this takes enormous effort on the part of both the competitors and the organizers. The Kaggle platform gives us a unique opportunity to reach many more people, and although there's a risk that the winning team ends up not generalizing very well, we think we've built enough safeguards into this challenge to minimize that risk.",
    "3305123": "gdclifford \n\n> In addition, we require that the teams' algorithms are fully retrainable, and we retrain them with a strict limit on compute time\n\nJust a confirmation, does this Kaggle competition has specific rule to restrict training time?",
    "3305213": "Training on massive data corpora has been an issue since around 2023, following a series of LLM competitions. But that’s fine — there are always some competitions that don’t require extensive GPU resources\n\nRecently, there was a Yale competition with a similar 'problem' (with the only diff that the tool wasn't provided, but it was relatively easy to reverse-engineer it) where teams with modest compute budgets (under $500) still managed to get gold. Of course, having GPUs is an advantage, but really smart and creative people can still win with ease",
    "3305269": "No - I'm referring to our other challenges where we have restricted both training and inference time. Kaggle competitions only restrict inference time per: [https://www.kaggle.com/docs/notebooks](https://www.kaggle.com/docs/notebooks) (look up \"Technical Specifications\")",
    "3305275": "10000TB data incoming",
    "3305282": "ecg-image-kit is not a synthetic ECG generator. It prints real input ECG time-series on realistic ECG grids, just like what a standard ECG machine does. The synthetic features of ecg-image-kit are the imaging artifacts and variations like rotation and wrinkles, not the ECG itself. The shared data have all been printed out in hardcopy and distorted with real world imaging artifacts. You can read more about the dataset in the references we've shared under the challenge description.",
    "3305285": "We did not generate artificial artifacts in the training or test images using ecg-image-kit. The images in the training and test data are created in the same way - we physically print real ECGs to ECG graph paper, then scan them back in (or photograph them) per the competition description. We also add real physical artifacts (creases, writing, stains, mold, etc.) and scan them. We provide ecg-image-kit to give you a way to simulate these artifacts, but we don't guarantee that it is exactly the same as the real ECG scans. It's up to you to judge how useful this is (perhaps by cross validating on the training data), and how much external (public?) data you need to use to generate enough useful training. You can also generate some images manually, like we did, if you do not think the artifacts generated are useful enough, or use some interesting training approaches that don't overfit on the artificial data you generate.",
    "3305306": "> The images in the training and test data are created in the same way - we physically print real ECGs to ECG graph paper, then scan them back in (or photograph them)\n\nAh, thanks for clarifying. This answers my question.\n\nI thought all the training images (including the augmented versions) came from ecg-image-kit because the data documentation says \"*train/[id]/[id]-0001.png Original color ECG image generated by ECG-image-kit*.\" However, with the additional information from your response, my current understanding is that you *only* used ecg-image-kit to generate the *clean* images; the augmented versions were created by physically printing and taking pictures of the clean images instead of using ecg-image-kit's built-in data augmentation functionality to simulate factors like wrinkles in the paper or rotation.",
    "3305317": "That is correct.",
    "3305326": "That's good new for me. I bought an 8TB SSD this year for kaggle; I only need to buy 1249 more!",
    "3305328": "should sponsor me some as well",
    "3305362": "I also bought 8TB SSD. 8TB SDD has 4800TBW, so you only need 16TB SSD in the best case.",
    "3305548": "Hello,\nI've recently completed my master's in data science and learnt aout AI/Ml.Also,I'm curious to know about the competition becuase it is my first time to participate can somebody help me with the rules and process where I can get an idea to proceed further.Thanks",
    "3305679": "And take 10000 TB of photos of moldy papers, it sure sounds realistic.",
    "3305802": "It depends on your model architecture. Do not underestimate using, or combining your methods with, more classical signal processing and computer vision techniques. Reading and incorporating contextual knowledge regarding the ECG also goes a long way.",
    "3306145": "Might can make a pure keypoint+alignment based method. I found that basic Orb approach can generate a lot of good points on the ecg image, then align each point with the actual timestamp in csv file would be nice. But if it's just for winning, the above 10000TB guy is still in favor of this competition.",
    "3307092": "But honestly, I’m curious to see if someone manages to pull a gold with clever signal processing instead of 10,000TB of synthetic ECGs.",
    "3307106": "People are begining to forget that $500 isn't that cheap...",
    "3308138": "true that just found this out the hard way lol"
  },
  "source": "meta"
}