{
  "id": 650669,
  "title": "Question about test images",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/650669",
  "author_name": "",
  "post_date": "2025-12-02T03:56:57.137581Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>What type of images in test set?</p>\n<p>I see that 2 example images in test set look like the train/[id]/[id]-0001.png Original color ECG image generated by ECG-image-kit.</p>\n<p>Do all test images look like this? Or it will be one of the types in train set</p>",
  "messages": [
    {
      "id": "3360322",
      "postDate": "12/02/2025 03:56:57",
      "content": "<p>What type of images in test set?</p>\n<p>I see that 2 example images in test set look like the train/[id]/[id]-0001.png Original color ECG image generated by ECG-image-kit.</p>\n<p>Do all test images look like this? Or it will be one of the types in train set</p>",
      "rawMarkdown": "What type of images in test set?\n\nI see that 2 example images in test set look like the train/[id]/[id]-0001.png Original color ECG image generated by ECG-image-kit.\n\nDo all test images look like this? Or it will be one of the types in train set",
      "votes": null
    },
    {
      "id": "3361265",
      "postDate": "12/03/2025 08:03:53",
      "content": "<p>I would checkout this discussion here:\n<a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729\" target=\"_blank\">https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729</a></p>\n<p>Specifically this response from <a href=\"https://www.kaggle.com/gdclifford\" target=\"_blank\">@gdclifford</a>:</p>\n<blockquote>\n  <p>\"This is a difficult question to answer, but thanks for posing it. We can state that all the data were generated using the same (non-synthetic) process (+/- some human variation in printing, scanning, and adding artifacts), and selected at random into the three sets (public and private (the 20:80 split for the leaderboard)). As a result, there are lots of visual similarities between the training and test data (both that used for the leaderboard and that used for the final score) but we can’t say they are ‘drawn from the same distribution’ or that the distribution is 'uniform' since that depends on what distribution/features you choose to measure and what statistical test you use to measure a meaningful difference in distributions. Therefore, we did not test the null hypotheses that any specific features were different between the distributions. We believe this is fine, because in reality, you always encounter samples that are out of the distribution of your training data, and it's part of the real-world challenge for you to think about ways to deal with this. This means that it may be really helpful if you build in some logic, rules and domain knowledge. We’ve provided lots of information in the comments to help you think of some ways to do this.\"</p>\n</blockquote>",
      "rawMarkdown": "I would checkout this discussion here:\n[https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729](https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729)\n\nSpecifically this response from @gdclifford:\n>\"This is a difficult question to answer, but thanks for posing it. We can state that all the data were generated using the same (non-synthetic) process (+/- some human variation in printing, scanning, and adding artifacts), and selected at random into the three sets (public and private (the 20:80 split for the leaderboard)). As a result, there are lots of visual similarities between the training and test data (both that used for the leaderboard and that used for the final score) but we can’t say they are ‘drawn from the same distribution’ or that the distribution is 'uniform' since that depends on what distribution/features you choose to measure and what statistical test you use to measure a meaningful difference in distributions. Therefore, we did not test the null hypotheses that any specific features were different between the distributions. We believe this is fine, because in reality, you always encounter samples that are out of the distribution of your training data, and it's part of the real-world challenge for you to think about ways to deal with this. This means that it may be really helpful if you build in some logic, rules and domain knowledge. We’ve provided lots of information in the comments to help you think of some ways to do this.\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3361265,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "12/03/2025 08:03:53",
      "content": "<p>I would checkout this discussion here:\n<a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729\" target=\"_blank\">https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729</a></p>\n<p>Specifically this response from <a href=\"https://www.kaggle.com/gdclifford\" target=\"_blank\">@gdclifford</a>:</p>\n<blockquote>\n  <p>\"This is a difficult question to answer, but thanks for posing it. We can state that all the data were generated using the same (non-synthetic) process (+/- some human variation in printing, scanning, and adding artifacts), and selected at random into the three sets (public and private (the 20:80 split for the leaderboard)). As a result, there are lots of visual similarities between the training and test data (both that used for the leaderboard and that used for the final score) but we can’t say they are ‘drawn from the same distribution’ or that the distribution is 'uniform' since that depends on what distribution/features you choose to measure and what statistical test you use to measure a meaningful difference in distributions. Therefore, we did not test the null hypotheses that any specific features were different between the distributions. We believe this is fine, because in reality, you always encounter samples that are out of the distribution of your training data, and it's part of the real-world challenge for you to think about ways to deal with this. This means that it may be really helpful if you build in some logic, rules and domain knowledge. We’ve provided lots of information in the comments to help you think of some ways to do this.\"</p>\n</blockquote>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3360322": "What type of images in test set?\n\nI see that 2 example images in test set look like the train/[id]/[id]-0001.png Original color ECG image generated by ECG-image-kit.\n\nDo all test images look like this? Or it will be one of the types in train set",
    "3361265": "I would checkout this discussion here:\n[https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729](https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729)\n\nSpecifically this response from @gdclifford:\n>\"This is a difficult question to answer, but thanks for posing it. We can state that all the data were generated using the same (non-synthetic) process (+/- some human variation in printing, scanning, and adding artifacts), and selected at random into the three sets (public and private (the 20:80 split for the leaderboard)). As a result, there are lots of visual similarities between the training and test data (both that used for the leaderboard and that used for the final score) but we can’t say they are ‘drawn from the same distribution’ or that the distribution is 'uniform' since that depends on what distribution/features you choose to measure and what statistical test you use to measure a meaningful difference in distributions. Therefore, we did not test the null hypotheses that any specific features were different between the distributions. We believe this is fine, because in reality, you always encounter samples that are out of the distribution of your training data, and it's part of the real-world challenge for you to think about ways to deal with this. This means that it may be really helpful if you build in some logic, rules and domain knowledge. We’ve provided lots of information in the comments to help you think of some ways to do this.\""
  },
  "source": "meta"
}