{
  "id": 613736,
  "title": "Converting from training csv files to data frame compatible with score function.",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/613736",
  "author_name": "",
  "post_date": "2025-10-29T10:57:29.440691800Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Does someone have a script to convert Pandas data frame <code>df</code> created from training csv file, e.g.</p>\n<pre><code>df  pd.read_csv'/kaggle/input/physionet-ecg-image-digitization/train/3/3.csv'\n</code></pre>\n<p>to the format required by <code>score</code> function from <a href=\"https://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric/\" target=\"_blank\">https://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric/</a> ?</p>",
  "messages": [
    {
      "id": "3308403",
      "postDate": "10/29/2025 10:57:29",
      "content": "<p>Does someone have a script to convert Pandas data frame <code>df</code> created from training csv file, e.g.</p>\n<pre><code>df  pd.read_csv'/kaggle/input/physionet-ecg-image-digitization/train/3/3.csv'\n</code></pre>\n<p>to the format required by <code>score</code> function from <a href=\"https://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric/\" target=\"_blank\">https://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric/</a> ?</p>",
      "rawMarkdown": "Does someone have a script to convert Pandas data frame `df` created from training csv file, e.g.\n```\ndf = pd.read_csv('/kaggle/input/physionet-ecg-image-digitization/train/7663343/7663343.csv')\n```\nto the format required by `score` function from https://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric/ ?",
      "votes": null
    },
    {
      "id": "3308625",
      "postDate": "10/29/2025 19:45:08",
      "content": "<blockquote>\n  <p>row_id_column_name = \"id\"</p>\n  <p>solution = pd.DataFrame({'id': ['343_0_I', '343_1_I', '343_2_I', '343_0_III', '343_1_III','343_2_III','343_0_aVR', '343_1_aVR','343_2_aVR',\\\n      '343_0_aVL', '343_1_aVL', '343_2_aVL', '343_0_aVF', '343_1_aVF','343_2_aVF','343_0_V1', '343_1_V1', '343_2_V1','343_0_V2', '343_1_V2','343_2_V2',\\\n      '343_0_V3', '343_1_V3', '343_2_V3','343_0_V4', '343_1_V4', '343_2_V4', '343_0_V5', '343_1_V5','343_2_V5','343_0_V6', '343_1_V6','343_2_V6',\\\n      '343_0_II', '343_1_II','343_2_II', '343_3_II', '343_4_II', '343_5_II','343_6_II', '343_7_II','343_8_II','343_9_II','343_10_II','343_11_II'],\\\n      'fs': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1],\\\n      'value':[0.1,0.3,0.4,0.6,0.6,0.4,0.2,0.3,0.4,0.5,0.2,0.7,0.2,0.3,0.4,0.8,0.6,0.7, 0.2,0.3,-0.1,0.5,0.6,0.7,0.2,0.9,0.4,0.5,0.6,0.7,0.1,0.3,0.4,\\\n      0.6,0.6,0.4,0.2,0.3,0.4,0.5,0.2,0.7,0.2,0.3,0.4]})</p>\n</blockquote>\n<p>EDIT: <a href=\"https://www.kaggle.com/code/sacuscreed/metric?scriptVersionId=271943371\" target=\"_blank\">https://www.kaggle.com/code/sacuscreed/metric?scriptVersionId=271943371</a></p>",
      "rawMarkdown": ">row_id_column_name = \"id\"\n\n>solution = pd.DataFrame({'id': ['343_0_I', '343_1_I', '343_2_I', '343_0_III', '343_1_III','343_2_III','343_0_aVR', '343_1_aVR','343_2_aVR',\\\n    '343_0_aVL', '343_1_aVL', '343_2_aVL', '343_0_aVF', '343_1_aVF','343_2_aVF','343_0_V1', '343_1_V1', '343_2_V1','343_0_V2', '343_1_V2','343_2_V2',\\\n    '343_0_V3', '343_1_V3', '343_2_V3','343_0_V4', '343_1_V4', '343_2_V4', '343_0_V5', '343_1_V5','343_2_V5','343_0_V6', '343_1_V6','343_2_V6',\\\n    '343_0_II', '343_1_II','343_2_II', '343_3_II', '343_4_II', '343_5_II','343_6_II', '343_7_II','343_8_II','343_9_II','343_10_II','343_11_II'],\\\n    'fs': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1],\\\n    'value':[0.1,0.3,0.4,0.6,0.6,0.4,0.2,0.3,0.4,0.5,0.2,0.7,0.2,0.3,0.4,0.8,0.6,0.7, 0.2,0.3,-0.1,0.5,0.6,0.7,0.2,0.9,0.4,0.5,0.6,0.7,0.1,0.3,0.4,\\\n    0.6,0.6,0.4,0.2,0.3,0.4,0.5,0.2,0.7,0.2,0.3,0.4]})\n\nEDIT: https://www.kaggle.com/code/sacuscreed/metric?scriptVersionId=271943371",
      "votes": null
    },
    {
      "id": "3308703",
      "postDate": "10/30/2025 04:40:51",
      "content": "<p>Thank you. Is there an evaluation function, which doesn't need the redundant <code>fs</code> column? There is an unnecessary proliferation of data formats in this competition:</p>\n<ol>\n<li><code>train/7663343/7663343.csv</code> format with <code>I,II,III,aVR,aVL,aVF,V1,V2,V3,V4,V5,V6</code> columns.</li>\n<li>Submission format with <code>id, value</code> columns.</li>\n<li>Evaluation format with <code>id, fs, value</code> columns.</li>\n</ol>\n<p>I wonder why? It seems that using only the first format would suffice.</p>",
      "rawMarkdown": "Thank you. Is there an evaluation function, which doesn't need the redundant `fs` column? There is an unnecessary proliferation of data formats in this competition:\n1. `train/7663343/7663343.csv` format with `I,II,III,aVR,aVL,aVF,V1,V2,V3,V4,V5,V6` columns.\n2. Submission format with `id, value` columns.\n3. Evaluation format with `id, fs, value` columns.\n\nI wonder why? It seems that using only the first format would suffice.",
      "votes": null
    },
    {
      "id": "3308818",
      "postDate": "10/30/2025 10:40:30",
      "content": "<p>That's only for local testing. Submission only needs id and values.</p>\n<blockquote>\n  <p>id,value</p>\n  <p>'62_0_I',0.0</p>\n  <p>'62_1_II',0.3</p>\n  <p>'62_2_I',0.4</p>\n  <p>etc.</p>\n</blockquote>",
      "rawMarkdown": "That's only for local testing. Submission only needs id and values.\n\n>id,value\n\n>'62_0_I',0.0\n\n>'62_1_II',0.3\n\n>'62_2_I',0.4\n\n>etc.",
      "votes": null
    },
    {
      "id": "3308899",
      "postDate": "10/30/2025 14:00:44",
      "content": "<p><code>fs</code> is not redundant. Please read <a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/613179#3306633\" target=\"_blank\">this</a> post.</p>",
      "rawMarkdown": "`fs` is not redundant. Please read [this](https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/613179#3306633) post.",
      "votes": null
    },
    {
      "id": "3309087",
      "postDate": "10/30/2025 23:17:21",
      "content": "<p>You are correct in general case, but it seems that in this competition all recordings are of 3x4+1 format and 10s long.</p>\n<blockquote>\n  <p>The expected number of rows is floor(fs * 10s) for lead II and floor(fs * 2.5s) for all other leads, where fs is the sampling frequency.</p>\n</blockquote>\n<p>Therefore, <code>fs</code> can be calculated as <code>ns/10</code> for lead II and <code>ns/2.5</code> for all other leads, where <code>ns</code> is a number of samples submitted for a given lead.</p>\n<p>Another way to prove my claim is the fact that you don't require <code>fs</code> column in submission file. It can be inferred using a method described above.</p>",
      "rawMarkdown": "You are correct in general case, but it seems that in this competition all recordings are of 3x4+1 format and 10s long.\n>The expected number of rows is floor(fs * 10s) for lead II and floor(fs * 2.5s) for all other leads, where fs is the sampling frequency.\n\nTherefore, `fs` can be calculated as `ns/10` for lead II and `ns/2.5` for all other leads, where `ns` is a number of samples submitted for a given lead.\n\nAnother way to prove my claim is the fact that you don't require `fs` column in submission file. It can be inferred using a method described above.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3308625,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "10/29/2025 19:45:08",
      "content": "<blockquote>\n  <p>row_id_column_name = \"id\"</p>\n  <p>solution = pd.DataFrame({'id': ['343_0_I', '343_1_I', '343_2_I', '343_0_III', '343_1_III','343_2_III','343_0_aVR', '343_1_aVR','343_2_aVR',\\\n      '343_0_aVL', '343_1_aVL', '343_2_aVL', '343_0_aVF', '343_1_aVF','343_2_aVF','343_0_V1', '343_1_V1', '343_2_V1','343_0_V2', '343_1_V2','343_2_V2',\\\n      '343_0_V3', '343_1_V3', '343_2_V3','343_0_V4', '343_1_V4', '343_2_V4', '343_0_V5', '343_1_V5','343_2_V5','343_0_V6', '343_1_V6','343_2_V6',\\\n      '343_0_II', '343_1_II','343_2_II', '343_3_II', '343_4_II', '343_5_II','343_6_II', '343_7_II','343_8_II','343_9_II','343_10_II','343_11_II'],\\\n      'fs': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1],\\\n      'value':[0.1,0.3,0.4,0.6,0.6,0.4,0.2,0.3,0.4,0.5,0.2,0.7,0.2,0.3,0.4,0.8,0.6,0.7, 0.2,0.3,-0.1,0.5,0.6,0.7,0.2,0.9,0.4,0.5,0.6,0.7,0.1,0.3,0.4,\\\n      0.6,0.6,0.4,0.2,0.3,0.4,0.5,0.2,0.7,0.2,0.3,0.4]})</p>\n</blockquote>\n<p>EDIT: <a href=\"https://www.kaggle.com/code/sacuscreed/metric?scriptVersionId=271943371\" target=\"_blank\">https://www.kaggle.com/code/sacuscreed/metric?scriptVersionId=271943371</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3308703,
          "author_name": "pauljurczak",
          "author_url": "",
          "post_date": "10/30/2025 04:40:51",
          "content": "<p>Thank you. Is there an evaluation function, which doesn't need the redundant <code>fs</code> column? There is an unnecessary proliferation of data formats in this competition:</p>\n<ol>\n<li><code>train/7663343/7663343.csv</code> format with <code>I,II,III,aVR,aVL,aVF,V1,V2,V3,V4,V5,V6</code> columns.</li>\n<li>Submission format with <code>id, value</code> columns.</li>\n<li>Evaluation format with <code>id, fs, value</code> columns.</li>\n</ol>\n<p>I wonder why? It seems that using only the first format would suffice.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3308818,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "10/30/2025 10:40:30",
              "content": "<p>That's only for local testing. Submission only needs id and values.</p>\n<blockquote>\n  <p>id,value</p>\n  <p>'62_0_I',0.0</p>\n  <p>'62_1_II',0.3</p>\n  <p>'62_2_I',0.4</p>\n  <p>etc.</p>\n</blockquote>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3308899,
              "author_name": "r2241272",
              "author_url": "",
              "post_date": "10/30/2025 14:00:44",
              "content": "<p><code>fs</code> is not redundant. Please read <a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/613179#3306633\" target=\"_blank\">this</a> post.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3309087,
                  "author_name": "pauljurczak",
                  "author_url": "",
                  "post_date": "10/30/2025 23:17:21",
                  "content": "<p>You are correct in general case, but it seems that in this competition all recordings are of 3x4+1 format and 10s long.</p>\n<blockquote>\n  <p>The expected number of rows is floor(fs * 10s) for lead II and floor(fs * 2.5s) for all other leads, where fs is the sampling frequency.</p>\n</blockquote>\n<p>Therefore, <code>fs</code> can be calculated as <code>ns/10</code> for lead II and <code>ns/2.5</code> for all other leads, where <code>ns</code> is a number of samples submitted for a given lead.</p>\n<p>Another way to prove my claim is the fact that you don't require <code>fs</code> column in submission file. It can be inferred using a method described above.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3308403": "Does someone have a script to convert Pandas data frame `df` created from training csv file, e.g.\n```\ndf = pd.read_csv('/kaggle/input/physionet-ecg-image-digitization/train/7663343/7663343.csv')\n```\nto the format required by `score` function from https://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric/ ?",
    "3308625": ">row_id_column_name = \"id\"\n\n>solution = pd.DataFrame({'id': ['343_0_I', '343_1_I', '343_2_I', '343_0_III', '343_1_III','343_2_III','343_0_aVR', '343_1_aVR','343_2_aVR',\\\n    '343_0_aVL', '343_1_aVL', '343_2_aVL', '343_0_aVF', '343_1_aVF','343_2_aVF','343_0_V1', '343_1_V1', '343_2_V1','343_0_V2', '343_1_V2','343_2_V2',\\\n    '343_0_V3', '343_1_V3', '343_2_V3','343_0_V4', '343_1_V4', '343_2_V4', '343_0_V5', '343_1_V5','343_2_V5','343_0_V6', '343_1_V6','343_2_V6',\\\n    '343_0_II', '343_1_II','343_2_II', '343_3_II', '343_4_II', '343_5_II','343_6_II', '343_7_II','343_8_II','343_9_II','343_10_II','343_11_II'],\\\n    'fs': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1],\\\n    'value':[0.1,0.3,0.4,0.6,0.6,0.4,0.2,0.3,0.4,0.5,0.2,0.7,0.2,0.3,0.4,0.8,0.6,0.7, 0.2,0.3,-0.1,0.5,0.6,0.7,0.2,0.9,0.4,0.5,0.6,0.7,0.1,0.3,0.4,\\\n    0.6,0.6,0.4,0.2,0.3,0.4,0.5,0.2,0.7,0.2,0.3,0.4]})\n\nEDIT: https://www.kaggle.com/code/sacuscreed/metric?scriptVersionId=271943371",
    "3308703": "Thank you. Is there an evaluation function, which doesn't need the redundant `fs` column? There is an unnecessary proliferation of data formats in this competition:\n1. `train/7663343/7663343.csv` format with `I,II,III,aVR,aVL,aVF,V1,V2,V3,V4,V5,V6` columns.\n2. Submission format with `id, value` columns.\n3. Evaluation format with `id, fs, value` columns.\n\nI wonder why? It seems that using only the first format would suffice.",
    "3308818": "That's only for local testing. Submission only needs id and values.\n\n>id,value\n\n>'62_0_I',0.0\n\n>'62_1_II',0.3\n\n>'62_2_I',0.4\n\n>etc.",
    "3308899": "`fs` is not redundant. Please read [this](https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/613179#3306633) post.",
    "3309087": "You are correct in general case, but it seems that in this competition all recordings are of 3x4+1 format and 10s long.\n>The expected number of rows is floor(fs * 10s) for lead II and floor(fs * 2.5s) for all other leads, where fs is the sampling frequency.\n\nTherefore, `fs` can be calculated as `ns/10` for lead II and `ns/2.5` for all other leads, where `ns` is a number of samples submitted for a given lead.\n\nAnother way to prove my claim is the fact that you don't require `fs` column in submission file. It can be inferred using a method described above."
  },
  "source": "meta"
}