{
  "id": 612929,
  "title": "At least 75% null values in 11/12 Leads? ",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/612929",
  "author_name": "samu2505",
  "post_date": "2025-10-22T23:46:28.275000",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>How to deal with the issue of huge missing values in the leads? Or is it part of the challenge that needs to be overcome? </p>\n<pre><code>LENGTHS = []\nNANS = []\n i, row  tqdm(train_df.iterrows()):\n    id_ = row[]\n    ecg_path = \n    ecg_df = pd.read_csv(ecg_path)\n    LENGTHS.append((ecg_df))\n    NANS.append((ecg_df.isna().() / (ecg_df)).values)\n</code></pre>\n<pre><code>np.stack(NANS).mean(axis=)\n</code></pre>\n<pre><code>array([,         , , , ,\n       , , , , ,\n       , ])\n</code></pre>",
  "messages": [
    {
      "id": 3305552,
      "postDate": "2025-10-23T00:09:08.717Z",
      "content": "<p>Yes - this is part of the challenge. It will be useful to take some time to understand what a 12 lead ECG is and how is laid out.  Only 2.5s of each 10s lead is plotted except for the rhythm lead, which is plotted in its entirety. I addressed that earlier today in this post:<br>\n<a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3305463\" target=\"_blank\">https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3305463</a><br>\nalthough I may not have been as clear as I'd hoped. The entries are only evaluated over non null values, so you only need to reconstruct the parts of the signal that you can see in the image. </p>",
      "rawMarkdown": "Yes - this is part of the challenge. It will be useful to take some time to understand what a 12 lead ECG is and how is laid out.  Only 2.5s of each 10s lead is plotted except for the rhythm lead, which is plotted in its entirety. I addressed that earlier today in this post:\nhttps://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3305463\nalthough I may not have been as clear as I'd hoped. The entries are only evaluated over non null values, so you only need to reconstruct the parts of the signal that you can see in the image. \n",
      "votes": 3,
      "replies": [
        {
          "id": 3308388,
          "postDate": "2025-10-29T10:00:57.837Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3305549,
      "postDate": "2025-10-22T23:46:28.277Z",
      "content": "<p>How to deal with the issue of huge missing values in the leads? Or is it part of the challenge that needs to be overcome? </p>\n<pre><code>LENGTHS = []\nNANS = []\n i, row  tqdm(train_df.iterrows()):\n    id_ = row[]\n    ecg_path = \n    ecg_df = pd.read_csv(ecg_path)\n    LENGTHS.append((ecg_df))\n    NANS.append((ecg_df.isna().() / (ecg_df)).values)\n</code></pre>\n<pre><code>np.stack(NANS).mean(axis=)\n</code></pre>\n<pre><code>array([,         , , , ,\n       , , , , ,\n       , ])\n</code></pre>",
      "rawMarkdown": "How to deal with the issue of huge missing values in the leads? Or is it part of the challenge that needs to be overcome? \n\n```python\nLENGTHS = []\nNANS = []\nfor i, row in tqdm(train_df.iterrows()):\n    id_ = row['id']\n    ecg_path = f\"{cfg.TRAIN_DATA_PATH}/{id_}/{id_}.csv\"\n    ecg_df = pd.read_csv(ecg_path)\n    LENGTHS.append(len(ecg_df))\n    NANS.append((ecg_df.isna().sum() / len(ecg_df)).values)\n```\n\n```python\nnp.stack(NANS).mean(axis=0)\n```\n\n```python\narray([0.75000809, 0.        , 0.75000809, 0.74999191, 0.74999191,\n       0.74999191, 0.75000809, 0.75000809, 0.75000809, 0.74999191,\n       0.74999191, 0.74999191])\n```",
      "votes": 3
    },
    {
      "id": 3305554,
      "postDate": "2025-10-23T00:16:40.853Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3305552,
      "author_name": "GDClifford",
      "author_url": "",
      "post_date": "2025-10-23T00:09:08.717000",
      "content": "<p>Yes - this is part of the challenge. It will be useful to take some time to understand what a 12 lead ECG is and how is laid out.  Only 2.5s of each 10s lead is plotted except for the rhythm lead, which is plotted in its entirety. I addressed that earlier today in this post:<br>\n<a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3305463\" target=\"_blank\">https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3305463</a><br>\nalthough I may not have been as clear as I'd hoped. The entries are only evaluated over non null values, so you only need to reconstruct the parts of the signal that you can see in the image. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 3308388,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-10-29T10:00:57.837000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3305554,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-10-23T00:16:40.853000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3305552": "Yes - this is part of the challenge. It will be useful to take some time to understand what a 12 lead ECG is and how is laid out.  Only 2.5s of each 10s lead is plotted except for the rhythm lead, which is plotted in its entirety. I addressed that earlier today in this post:\nhttps://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/612729#3305463\nalthough I may not have been as clear as I'd hoped. The entries are only evaluated over non null values, so you only need to reconstruct the parts of the signal that you can see in the image. \n",
    "3305549": "How to deal with the issue of huge missing values in the leads? Or is it part of the challenge that needs to be overcome? \n\n```python\nLENGTHS = []\nNANS = []\nfor i, row in tqdm(train_df.iterrows()):\n    id_ = row['id']\n    ecg_path = f\"{cfg.TRAIN_DATA_PATH}/{id_}/{id_}.csv\"\n    ecg_df = pd.read_csv(ecg_path)\n    LENGTHS.append(len(ecg_df))\n    NANS.append((ecg_df.isna().sum() / len(ecg_df)).values)\n```\n\n```python\nnp.stack(NANS).mean(axis=0)\n```\n\n```python\narray([0.75000809, 0.        , 0.75000809, 0.74999191, 0.74999191,\n       0.74999191, 0.75000809, 0.75000809, 0.75000809, 0.74999191,\n       0.74999191, 0.74999191])\n```",
    "3305554": ""
  }
}