{
  "id": 613179,
  "title": "Why do we have sample rate mentioned in test.csv when it is not needed?",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/613179",
  "author_name": "",
  "post_date": "2025-10-24T19:39:15.978574500Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Will the sample rate be of any use anywhere?</p>",
  "messages": [
    {
      "id": "3306570",
      "postDate": "10/24/2025 19:39:15",
      "content": "<p>Will the sample rate be of any use anywhere?</p>",
      "rawMarkdown": "Will the sample rate be of any use anywhere?",
      "votes": null
    },
    {
      "id": "3306633",
      "postDate": "10/25/2025 00:25:58",
      "content": "<p>We have addressed this in detail in the Appendix of the <a href=\"https://doi.org/10.48550/arXiv.2409.16612\" target=\"_blank\">ECG-Image-Dataset</a> preprint. In summary, the reference ECG signal has a sampling frequency of, say, <code>fs</code>. When it is printed on paper or displayed on a screen, it is back to the analog domain; when it is scanned or photographed, it is converted back to the digital domain, this time with another sampling frequency of, say, <code>fs'</code>, which depends on the image resolution rather than the original sampling frequency <code>fs</code>. When you extract the time series from the images, you need to know the original time-series sampling frequency <code>fs</code> to be able to resample your estimated samples to that frequency and compare them for scoring and visualization. We have a similar discussion regarding signal amplitude resolution and scales, which you can follow in the <a href=\"https://doi.org/10.48550/arXiv.2409.16612\" target=\"_blank\">ECG-Image-Dataset</a> preprint and the other references we have shared.</p>",
      "rawMarkdown": "We have addressed this in detail in the Appendix of the [ECG-Image-Dataset](https://doi.org/10.48550/arXiv.2409.16612) preprint. In summary, the reference ECG signal has a sampling frequency of, say, `fs`. When it is printed on paper or displayed on a screen, it is back to the analog domain; when it is scanned or photographed, it is converted back to the digital domain, this time with another sampling frequency of, say, `fs'`, which depends on the image resolution rather than the original sampling frequency `fs`. When you extract the time series from the images, you need to know the original time-series sampling frequency `fs` to be able to resample your estimated samples to that frequency and compare them for scoring and visualization. We have a similar discussion regarding signal amplitude resolution and scales, which you can follow in the [ECG-Image-Dataset](https://doi.org/10.48550/arXiv.2409.16612) preprint and the other references we have shared.",
      "votes": null
    },
    {
      "id": "3306647",
      "postDate": "10/25/2025 02:21:39",
      "content": "<p><a href=\"https://www.kaggle.com/r2241272\" target=\"_blank\">@r2241272</a> is it safe to create N predictions using fs * 2.5 (except II, which is *10) or should we create N predictions using the number_of_rows? I got some failed submissions and im trying to narrow down possible issues.</p>\n<p>Also, i ignore if the row ids should be in a specific order.</p>\n<p>Should we check if:</p>\n<ul>\n<li>id is of type object</li>\n<li>value is of type float</li>\n<li>id is unique</li>\n<li>there are as many rows in the submission file as the total sum of number_of_rows column in the test csv</li>\n<li>no NaN values?</li>\n</ul>\n<p>I dont know if im missing anything else. Could you confirm this?</p>",
      "rawMarkdown": "r2241272 is it safe to create N predictions using fs * 2.5 (except II, which is *10) or should we create N predictions using the number_of_rows? I got some failed submissions and im trying to narrow down possible issues.\n\nAlso, i ignore if the row ids should be in a specific order.\n\n Should we check if:\n- id is of type object\n- value is of type float\n- id is unique\n- there are as many rows in the submission file as the total sum of number_of_rows column in the test csv\n- no NaN values?\n\nI dont know if im missing anything else. Could you confirm this?",
      "votes": null
    },
    {
      "id": "3306680",
      "postDate": "10/25/2025 03:21:29",
      "content": "<p>The evaluation metric code for the competition expects <code>number_of_rows</code> samples for each lead, which may or may not be the same as <code>fs * 2.5</code> or <code>fs * 10</code> for certain values of <code>fs</code>, <code>2.5</code>, and <code>10</code>.</p>\n<p>The evaluation metric code suggests some things to check, and, of course, the more robust your code, the better. For example, in general, not every signal is 10 seconds long, and code that is robust to variation outside of this competition may very well be better within the competition.</p>",
      "rawMarkdown": "The evaluation metric code for the competition expects `number_of_rows` samples for each lead, which may or may not be the same as `fs * 2.5` or `fs * 10` for certain values of `fs`, `2.5`, and `10`.\n\nThe evaluation metric code suggests some things to check, and, of course, the more robust your code, the better. For example, in general, not every signal is 10 seconds long, and code that is robust to variation outside of this competition may very well be better within the competition.",
      "votes": null
    },
    {
      "id": "3306701",
      "postDate": "10/25/2025 04:27:15",
      "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> For some images, fs is odd, and fs * 2.5 is not integer. In train, they sometimes round up and sometimes down; in test, you always have to round down.</p>\n<blockquote>\n  <p>The expected number of rows is floor(fs * 10s) for lead II and floor(fs * 2.5s) for all other leads.</p>\n</blockquote>",
      "rawMarkdown": "alejopaullier For some images, fs is odd, and fs \\* 2.5 is not integer. In train, they sometimes round up and sometimes down; in test, you always have to round down.\n\n> The expected number of rows is floor(fs * 10s) for lead II and floor(fs * 2.5s) for all other leads.",
      "votes": null
    },
    {
      "id": "3306867",
      "postDate": "10/25/2025 12:59:12",
      "content": "<p>As <a href=\"https://www.kaggle.com/matthewreyna\" target=\"_blank\">@matthewreyna</a> noted, read the sampling frequency and expected number of samples from <code>train.csv</code> and <code>test.csv</code> instead of hardcoding the signal lengths to ensure your code remains portable and generalizable. ECG data generated by different machines can have varying formats. For example, check out the introductory chapters on ECG formats of <a href=\"https://landing1.gehealthcare.com/rs/005-SHS-767/images/45351-MUSE-17Nov2022-6-1-Quick-Reference-Guide-LP-Diagnostic-Cardiology.pdf\" target=\"_blank\">GE</a> and <a href=\"https://www.documents.philips.com/assets/Instruction%20for%20Use/20250404/4dce4266a3d64303985db2b500dfb2dc.pdf\" target=\"_blank\">Philips</a> ECG machines to see how they sweep the leads over time in their printouts. The <code>NaN</code> values correspond to samples that are not printed on the ECG image.</p>",
      "rawMarkdown": "As @matthewreyna noted, read the sampling frequency and expected number of samples from `train.csv` and `test.csv` instead of hardcoding the signal lengths to ensure your code remains portable and generalizable. ECG data generated by different machines can have varying formats. For example, check out the introductory chapters on ECG formats of [GE](https://landing1.gehealthcare.com/rs/005-SHS-767/images/45351-MUSE-17Nov2022-6-1-Quick-Reference-Guide-LP-Diagnostic-Cardiology.pdf) and [Philips](https://www.documents.philips.com/assets/Instruction%20for%20Use/20250404/4dce4266a3d64303985db2b500dfb2dc.pdf) ECG machines to see how they sweep the leads over time in their printouts. The `NaN` values correspond to samples that are not printed on the ECG image.",
      "votes": null
    },
    {
      "id": "3306901",
      "postDate": "10/25/2025 14:43:25",
      "content": "<p>That's weird. I have now changed my code to use <code>number_of_rows</code> and I am still getting:</p>\n<pre><code>Your notebook  a submission file  incorrect .  examples causing this are: wrong number    , empty , an incorrect data   a ,  invalid submission   what  expected.\n</code></pre>\n<p>These assertions have been passed successfully without the notebook throwing an exception error:</p>\n<pre><code> submission[\"id\"].dtype == np.dtype(\"O\")  #   id   correct \n submission[\"value\"].dtype == np.dtype(\"float64\")  #      correct \n submission[\"id\"].nunique() == len(submission)  #   id  \n  submission[\"value\"].isna().()  #     \n len(submission) == df_test.number_of_rows.sum()  #   submission has correct number  \n</code></pre>",
      "rawMarkdown": "That's weird. I have now changed my code to use `number_of_rows` and I am still getting:\n```\nYour notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected.\n```\nThese assertions have been passed successfully without the notebook throwing an exception error:\n```\nassert submission[\"id\"].dtype == np.dtype(\"O\")  # check if id is of correct type\nassert submission[\"value\"].dtype == np.dtype(\"float64\")  # check if value is of correct type\nassert submission[\"id\"].nunique() == len(submission)  # check if id is unique\nassert not submission[\"value\"].isna().any()  # check for no nan values\nassert len(submission) == df_test.number_of_rows.sum()  # check if submission has correct number of rows\n```",
      "votes": null
    },
    {
      "id": "3306916",
      "postDate": "10/25/2025 15:23:39",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> could you confirm these assertions/checks are correct? I have not been able to determine the root cause of the issue</p>",
      "rawMarkdown": "sohier could you confirm these assertions/checks are correct? I have not been able to determine the root cause of the issue",
      "votes": null
    },
    {
      "id": "3307742",
      "postDate": "10/27/2025 17:11:06",
      "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>, I noticed here you are checking the sum of number_of_rows. Please check if the prediction length is correct for each individual lead.</p>",
      "rawMarkdown": "alejopaullier, I noticed here you are checking the sum of number_of_rows. Please check if the prediction length is correct for each individual lead.",
      "votes": null
    },
    {
      "id": "3312372",
      "postDate": "11/07/2025 04:30:10",
      "content": "<p>Have you found the cause of the problem? I am facing the same issue. From what I can tell:</p>\n<ol>\n<li>Number of rows is correct</li>\n<li>Number of columns is correct. Column names and order are correct.</li>\n<li>No NaN, empty, non-finite.</li>\n<li>Number of rows per lead follows text.csv num_of_rows.</li>\n<li>All exams include the 12 leads</li>\n</ol>\n<p>I am still getting: Submission Scoring Error</p>",
      "rawMarkdown": "Have you found the cause of the problem? I am facing the same issue. From what I can tell:\n1. Number of rows is correct\n2. Number of columns is correct. Column names and order are correct.\n3. No NaN, empty, non-finite.\n4. Number of rows per lead follows text.csv num_of_rows.\n5. All exams include the 12 leads\n\nI am still getting: Submission Scoring Error",
      "votes": null
    },
    {
      "id": "3312690",
      "postDate": "11/07/2025 18:06:39",
      "content": "<p><a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a> I do not know exactly why my old code was failing. To generate predictions I was looping through all unique <code>base_id</code> and grabbing the test dataframe subset corresponding to that <code>base_id</code> and generating predictions for that subset, then concatenate all subsets into <code>submission</code>. </p>\n<p>Then I decided to loop over test rows using <code>df_test.iterrows()</code>  and the problem went away. Perhaps there is something odd for some specific ids.</p>",
      "rawMarkdown": "felipekitamura I do not know exactly why my old code was failing. To generate predictions I was looping through all unique `base_id` and grabbing the test dataframe subset corresponding to that `base_id` and generating predictions for that subset, then concatenate all subsets into `submission`. \n\nThen I decided to loop over test rows using `df_test.iterrows()`  and the problem went away. Perhaps there is something odd for some specific ids.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3306633,
      "author_name": "r2241272",
      "author_url": "",
      "post_date": "10/25/2025 00:25:58",
      "content": "<p>We have addressed this in detail in the Appendix of the <a href=\"https://doi.org/10.48550/arXiv.2409.16612\" target=\"_blank\">ECG-Image-Dataset</a> preprint. In summary, the reference ECG signal has a sampling frequency of, say, <code>fs</code>. When it is printed on paper or displayed on a screen, it is back to the analog domain; when it is scanned or photographed, it is converted back to the digital domain, this time with another sampling frequency of, say, <code>fs'</code>, which depends on the image resolution rather than the original sampling frequency <code>fs</code>. When you extract the time series from the images, you need to know the original time-series sampling frequency <code>fs</code> to be able to resample your estimated samples to that frequency and compare them for scoring and visualization. We have a similar discussion regarding signal amplitude resolution and scales, which you can follow in the <a href=\"https://doi.org/10.48550/arXiv.2409.16612\" target=\"_blank\">ECG-Image-Dataset</a> preprint and the other references we have shared.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3306647,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "10/25/2025 02:21:39",
          "content": "<p><a href=\"https://www.kaggle.com/r2241272\" target=\"_blank\">@r2241272</a> is it safe to create N predictions using fs * 2.5 (except II, which is *10) or should we create N predictions using the number_of_rows? I got some failed submissions and im trying to narrow down possible issues.</p>\n<p>Also, i ignore if the row ids should be in a specific order.</p>\n<p>Should we check if:</p>\n<ul>\n<li>id is of type object</li>\n<li>value is of type float</li>\n<li>id is unique</li>\n<li>there are as many rows in the submission file as the total sum of number_of_rows column in the test csv</li>\n<li>no NaN values?</li>\n</ul>\n<p>I dont know if im missing anything else. Could you confirm this?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3306680,
              "author_name": "matthewreyna",
              "author_url": "",
              "post_date": "10/25/2025 03:21:29",
              "content": "<p>The evaluation metric code for the competition expects <code>number_of_rows</code> samples for each lead, which may or may not be the same as <code>fs * 2.5</code> or <code>fs * 10</code> for certain values of <code>fs</code>, <code>2.5</code>, and <code>10</code>.</p>\n<p>The evaluation metric code suggests some things to check, and, of course, the more robust your code, the better. For example, in general, not every signal is 10 seconds long, and code that is robust to variation outside of this competition may very well be better within the competition.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3306701,
              "author_name": "ambrosm",
              "author_url": "",
              "post_date": "10/25/2025 04:27:15",
              "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> For some images, fs is odd, and fs * 2.5 is not integer. In train, they sometimes round up and sometimes down; in test, you always have to round down.</p>\n<blockquote>\n  <p>The expected number of rows is floor(fs * 10s) for lead II and floor(fs * 2.5s) for all other leads.</p>\n</blockquote>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3306867,
              "author_name": "r2241272",
              "author_url": "",
              "post_date": "10/25/2025 12:59:12",
              "content": "<p>As <a href=\"https://www.kaggle.com/matthewreyna\" target=\"_blank\">@matthewreyna</a> noted, read the sampling frequency and expected number of samples from <code>train.csv</code> and <code>test.csv</code> instead of hardcoding the signal lengths to ensure your code remains portable and generalizable. ECG data generated by different machines can have varying formats. For example, check out the introductory chapters on ECG formats of <a href=\"https://landing1.gehealthcare.com/rs/005-SHS-767/images/45351-MUSE-17Nov2022-6-1-Quick-Reference-Guide-LP-Diagnostic-Cardiology.pdf\" target=\"_blank\">GE</a> and <a href=\"https://www.documents.philips.com/assets/Instruction%20for%20Use/20250404/4dce4266a3d64303985db2b500dfb2dc.pdf\" target=\"_blank\">Philips</a> ECG machines to see how they sweep the leads over time in their printouts. The <code>NaN</code> values correspond to samples that are not printed on the ECG image.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3306901,
              "author_name": "alejopaullier",
              "author_url": "",
              "post_date": "10/25/2025 14:43:25",
              "content": "<p>That's weird. I have now changed my code to use <code>number_of_rows</code> and I am still getting:</p>\n<pre><code>Your notebook  a submission file  incorrect .  examples causing this are: wrong number    , empty , an incorrect data   a ,  invalid submission   what  expected.\n</code></pre>\n<p>These assertions have been passed successfully without the notebook throwing an exception error:</p>\n<pre><code> submission[\"id\"].dtype == np.dtype(\"O\")  #   id   correct \n submission[\"value\"].dtype == np.dtype(\"float64\")  #      correct \n submission[\"id\"].nunique() == len(submission)  #   id  \n  submission[\"value\"].isna().()  #     \n len(submission) == df_test.number_of_rows.sum()  #   submission has correct number  \n</code></pre>",
              "votes": null,
              "replies": [
                {
                  "id": 3306916,
                  "author_name": "alejopaullier",
                  "author_url": "",
                  "post_date": "10/25/2025 15:23:39",
                  "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> could you confirm these assertions/checks are correct? I have not been able to determine the root cause of the issue</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3307742,
                      "author_name": "yanyao23",
                      "author_url": "",
                      "post_date": "10/27/2025 17:11:06",
                      "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>, I noticed here you are checking the sum of number_of_rows. Please check if the prediction length is correct for each individual lead.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            },
            {
              "id": 3312372,
              "author_name": "felipekitamura",
              "author_url": "",
              "post_date": "11/07/2025 04:30:10",
              "content": "<p>Have you found the cause of the problem? I am facing the same issue. From what I can tell:</p>\n<ol>\n<li>Number of rows is correct</li>\n<li>Number of columns is correct. Column names and order are correct.</li>\n<li>No NaN, empty, non-finite.</li>\n<li>Number of rows per lead follows text.csv num_of_rows.</li>\n<li>All exams include the 12 leads</li>\n</ol>\n<p>I am still getting: Submission Scoring Error</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3312690,
                  "author_name": "alejopaullier",
                  "author_url": "",
                  "post_date": "11/07/2025 18:06:39",
                  "content": "<p><a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a> I do not know exactly why my old code was failing. To generate predictions I was looping through all unique <code>base_id</code> and grabbing the test dataframe subset corresponding to that <code>base_id</code> and generating predictions for that subset, then concatenate all subsets into <code>submission</code>. </p>\n<p>Then I decided to loop over test rows using <code>df_test.iterrows()</code>  and the problem went away. Perhaps there is something odd for some specific ids.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3306570": "Will the sample rate be of any use anywhere?",
    "3306633": "We have addressed this in detail in the Appendix of the [ECG-Image-Dataset](https://doi.org/10.48550/arXiv.2409.16612) preprint. In summary, the reference ECG signal has a sampling frequency of, say, `fs`. When it is printed on paper or displayed on a screen, it is back to the analog domain; when it is scanned or photographed, it is converted back to the digital domain, this time with another sampling frequency of, say, `fs'`, which depends on the image resolution rather than the original sampling frequency `fs`. When you extract the time series from the images, you need to know the original time-series sampling frequency `fs` to be able to resample your estimated samples to that frequency and compare them for scoring and visualization. We have a similar discussion regarding signal amplitude resolution and scales, which you can follow in the [ECG-Image-Dataset](https://doi.org/10.48550/arXiv.2409.16612) preprint and the other references we have shared.",
    "3306647": "r2241272 is it safe to create N predictions using fs * 2.5 (except II, which is *10) or should we create N predictions using the number_of_rows? I got some failed submissions and im trying to narrow down possible issues.\n\nAlso, i ignore if the row ids should be in a specific order.\n\n Should we check if:\n- id is of type object\n- value is of type float\n- id is unique\n- there are as many rows in the submission file as the total sum of number_of_rows column in the test csv\n- no NaN values?\n\nI dont know if im missing anything else. Could you confirm this?",
    "3306680": "The evaluation metric code for the competition expects `number_of_rows` samples for each lead, which may or may not be the same as `fs * 2.5` or `fs * 10` for certain values of `fs`, `2.5`, and `10`.\n\nThe evaluation metric code suggests some things to check, and, of course, the more robust your code, the better. For example, in general, not every signal is 10 seconds long, and code that is robust to variation outside of this competition may very well be better within the competition.",
    "3306701": "alejopaullier For some images, fs is odd, and fs \\* 2.5 is not integer. In train, they sometimes round up and sometimes down; in test, you always have to round down.\n\n> The expected number of rows is floor(fs * 10s) for lead II and floor(fs * 2.5s) for all other leads.",
    "3306867": "As @matthewreyna noted, read the sampling frequency and expected number of samples from `train.csv` and `test.csv` instead of hardcoding the signal lengths to ensure your code remains portable and generalizable. ECG data generated by different machines can have varying formats. For example, check out the introductory chapters on ECG formats of [GE](https://landing1.gehealthcare.com/rs/005-SHS-767/images/45351-MUSE-17Nov2022-6-1-Quick-Reference-Guide-LP-Diagnostic-Cardiology.pdf) and [Philips](https://www.documents.philips.com/assets/Instruction%20for%20Use/20250404/4dce4266a3d64303985db2b500dfb2dc.pdf) ECG machines to see how they sweep the leads over time in their printouts. The `NaN` values correspond to samples that are not printed on the ECG image.",
    "3306901": "That's weird. I have now changed my code to use `number_of_rows` and I am still getting:\n```\nYour notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected.\n```\nThese assertions have been passed successfully without the notebook throwing an exception error:\n```\nassert submission[\"id\"].dtype == np.dtype(\"O\")  # check if id is of correct type\nassert submission[\"value\"].dtype == np.dtype(\"float64\")  # check if value is of correct type\nassert submission[\"id\"].nunique() == len(submission)  # check if id is unique\nassert not submission[\"value\"].isna().any()  # check for no nan values\nassert len(submission) == df_test.number_of_rows.sum()  # check if submission has correct number of rows\n```",
    "3306916": "sohier could you confirm these assertions/checks are correct? I have not been able to determine the root cause of the issue",
    "3307742": "alejopaullier, I noticed here you are checking the sum of number_of_rows. Please check if the prediction length is correct for each individual lead.",
    "3312372": "Have you found the cause of the problem? I am facing the same issue. From what I can tell:\n1. Number of rows is correct\n2. Number of columns is correct. Column names and order are correct.\n3. No NaN, empty, non-finite.\n4. Number of rows per lead follows text.csv num_of_rows.\n5. All exams include the 12 leads\n\nI am still getting: Submission Scoring Error",
    "3312690": "felipekitamura I do not know exactly why my old code was failing. To generate predictions I was looping through all unique `base_id` and grabbing the test dataframe subset corresponding to that `base_id` and generating predictions for that subset, then concatenate all subsets into `submission`. \n\nThen I decided to loop over test rows using `df_test.iterrows()`  and the problem went away. Perhaps there is something odd for some specific ids."
  },
  "source": "meta"
}