{
  "id": 613480,
  "title": "Help needed for a complete beginner",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/613480",
  "author_name": "",
  "post_date": "2025-10-27T06:25:58.493155800Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>Thanks for anyone's help and contribution in advance.</p>\n<p>Basically, I'm almost like a complete beginner in this field, Signal Processing or ECG, and there are a couple of questions I'd like to ask to the community to understand the goal of the competition more accurately.</p>\n<p><strong>1. What does the reconstruction of ECG to a time-series data mean?</strong><br>\nI think the dataset has the original printout of ECG, and their time-series values in csv file, if I've understood it right. For example, 1000000 folder would contain 1000000-0001.png and 1000000.csv. Does the competition want us to build a model that could regressively output the time-series data based on the various types of training images?</p>\n<p><strong>2. How is the submission file's <code>value</code> column calculated from the reconstruction of the time-series data?</strong><br>\nThis might be related to the first question, but I can't see how the time-series reconstruction, which sounds very multi-dimensional, can be summed up to only one value per data.</p>\n<p><strong>Note</strong> that I've seen other notebooks and am trying to understand them through reading the book referred to by the host, so please do not just copy and paste a link that only contains the information I need at the end of its about 30-page long document.</p>",
  "messages": [
    {
      "id": "3307500",
      "postDate": "10/27/2025 06:25:58",
      "content": "<p>Hi,</p>\n<p>Thanks for anyone's help and contribution in advance.</p>\n<p>Basically, I'm almost like a complete beginner in this field, Signal Processing or ECG, and there are a couple of questions I'd like to ask to the community to understand the goal of the competition more accurately.</p>\n<p><strong>1. What does the reconstruction of ECG to a time-series data mean?</strong><br>\nI think the dataset has the original printout of ECG, and their time-series values in csv file, if I've understood it right. For example, 1000000 folder would contain 1000000-0001.png and 1000000.csv. Does the competition want us to build a model that could regressively output the time-series data based on the various types of training images?</p>\n<p><strong>2. How is the submission file's <code>value</code> column calculated from the reconstruction of the time-series data?</strong><br>\nThis might be related to the first question, but I can't see how the time-series reconstruction, which sounds very multi-dimensional, can be summed up to only one value per data.</p>\n<p><strong>Note</strong> that I've seen other notebooks and am trying to understand them through reading the book referred to by the host, so please do not just copy and paste a link that only contains the information I need at the end of its about 30-page long document.</p>",
      "rawMarkdown": "Hi,\n\nThanks for anyone's help and contribution in advance.\n\nBasically, I'm almost like a complete beginner in this field, Signal Processing or ECG, and there are a couple of questions I'd like to ask to the community to understand the goal of the competition more accurately.\n\n**1. What does the reconstruction of ECG to a time-series data mean?**\nI think the dataset has the original printout of ECG, and their time-series values in csv file, if I've understood it right. For example, 1000000 folder would contain 1000000-0001.png and 1000000.csv. Does the competition want us to build a model that could regressively output the time-series data based on the various types of training images?\n\n**2. How is the submission file's `value` column calculated from the reconstruction of the time-series data?**\nThis might be related to the first question, but I can't see how the time-series reconstruction, which sounds very multi-dimensional, can be summed up to only one value per data.\n\n\n**Note** that I've seen other notebooks and am trying to understand them through reading the book referred to by the host, so please do not just copy and paste a link that only contains the information I need at the end of its about 30-page long document.",
      "votes": null
    },
    {
      "id": "3308585",
      "postDate": "10/29/2025 17:47:25",
      "content": "<p>Time series data in each folder has come from real life sources. That is the core truth. ECG Image kit and other manual methods were then used to generate various images as seen in the folders. Check <a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3305038\" target=\"_blank\">https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3305038</a></p>\n<p>You need to train your model on the images in the folders to against the time series from that folder. Images are of different class where some has been soiled and some molded. Your model should expect to encounter any kind of such images when doing prediction. You will use train your model on images and validate results against the csv files in those folders. Build models against all kind of images. Predict X and Y values from images for different intervals and validate those against base truth in CSV files.</p>\n<p>Once model is built and ready, run it against the 2 test images, generate time series data for those, and submit that CSV.</p>\n<p>Your model would be executed against the ptivate test images and score would be assigned. Overview tab has details on time required. Check <a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3307329\" target=\"_blank\">https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3307329</a> for details of the private test images. Those private test images are not shared here, for obvius reasons.  </p>",
      "rawMarkdown": "Time series data in each folder has come from real life sources. That is the core truth. ECG Image kit and other manual methods were then used to generate various images as seen in the folders. Check https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3305038\n\nYou need to train your model on the images in the folders to against the time series from that folder. Images are of different class where some has been soiled and some molded. Your model should expect to encounter any kind of such images when doing prediction. You will use train your model on images and validate results against the csv files in those folders. Build models against all kind of images. Predict X and Y values from images for different intervals and validate those against base truth in CSV files.\n\nOnce model is built and ready, run it against the 2 test images, generate time series data for those, and submit that CSV.\n\nYour model would be executed against the ptivate test images and score would be assigned. Overview tab has details on time required. Check https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3307329 for details of the private test images. Those private test images are not shared here, for obvius reasons.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3308585,
      "author_name": "atulkumar1",
      "author_url": "",
      "post_date": "10/29/2025 17:47:25",
      "content": "<p>Time series data in each folder has come from real life sources. That is the core truth. ECG Image kit and other manual methods were then used to generate various images as seen in the folders. Check <a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3305038\" target=\"_blank\">https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3305038</a></p>\n<p>You need to train your model on the images in the folders to against the time series from that folder. Images are of different class where some has been soiled and some molded. Your model should expect to encounter any kind of such images when doing prediction. You will use train your model on images and validate results against the csv files in those folders. Build models against all kind of images. Predict X and Y values from images for different intervals and validate those against base truth in CSV files.</p>\n<p>Once model is built and ready, run it against the 2 test images, generate time series data for those, and submit that CSV.</p>\n<p>Your model would be executed against the ptivate test images and score would be assigned. Overview tab has details on time required. Check <a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3307329\" target=\"_blank\">https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3307329</a> for details of the private test images. Those private test images are not shared here, for obvius reasons.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3307500": "Hi,\n\nThanks for anyone's help and contribution in advance.\n\nBasically, I'm almost like a complete beginner in this field, Signal Processing or ECG, and there are a couple of questions I'd like to ask to the community to understand the goal of the competition more accurately.\n\n**1. What does the reconstruction of ECG to a time-series data mean?**\nI think the dataset has the original printout of ECG, and their time-series values in csv file, if I've understood it right. For example, 1000000 folder would contain 1000000-0001.png and 1000000.csv. Does the competition want us to build a model that could regressively output the time-series data based on the various types of training images?\n\n**2. How is the submission file's `value` column calculated from the reconstruction of the time-series data?**\nThis might be related to the first question, but I can't see how the time-series reconstruction, which sounds very multi-dimensional, can be summed up to only one value per data.\n\n\n**Note** that I've seen other notebooks and am trying to understand them through reading the book referred to by the host, so please do not just copy and paste a link that only contains the information I need at the end of its about 30-page long document.",
    "3308585": "Time series data in each folder has come from real life sources. That is the core truth. ECG Image kit and other manual methods were then used to generate various images as seen in the folders. Check https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3305038\n\nYou need to train your model on the images in the folders to against the time series from that folder. Images are of different class where some has been soiled and some molded. Your model should expect to encounter any kind of such images when doing prediction. You will use train your model on images and validate results against the csv files in those folders. Build models against all kind of images. Predict X and Y values from images for different intervals and validate those against base truth in CSV files.\n\nOnce model is built and ready, run it against the 2 test images, generate time series data for those, and submit that CSV.\n\nYour model would be executed against the ptivate test images and score would be assigned. Overview tab has details on time required. Check https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3307329 for details of the private test images. Those private test images are not shared here, for obvius reasons."
  },
  "source": "meta"
}