{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":97984,"databundleVersionId":14096757,"sourceType":"competition"}],"dockerImageVersionId":31234,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# ECG_Digitization_v4_Processing Pipeline (Centerline → Waveform)\n\nThis notebook focuses on post-processing improvements\nfor the v4 hybrid centerline extraction method.\n\nThe aim is to stabilize the waveform and improve\nsubmission score in the PhysioNet ECG digitization challenge.\n","metadata":{}},{"cell_type":"markdown","source":"## Overview\n\nThis notebook generates a submission CSV for the PhysioNet ECG Image Digitization challenge.\n\nKey idea:\n- Extract ECG centerline from grid-removed images\n- Normalize vertical position (pixel-based) into a stable waveform representation\n- Focus on robust shape preservation rather than absolute mV calibration\n\n","metadata":{}},{"cell_type":"markdown","source":"## Sample Output (v4: hybrid + smoothing)\n\nBelow is an example of the extracted ECG waveform from a single input image\nafter centerline extraction, hybrid post-processing, and smoothing.\n\n- Input: PhysioNet ECG image\n- Output: 1D waveform (normalized amplitude)\n","metadata":{}},{"cell_type":"markdown","source":"###  1. Imports & Environment","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport os, cv2, numpy as np, pandas as pd, matplotlib.pyplot as plt\n\nfrom scipy.ndimage import median_filter, gaussian_filter1d\nfrom scipy.signal import savgol_filter\n\nplt.rcParams[\"figure.figsize\"] = (12,4)\nplt.rcParams[\"axes.grid\"] = True\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-19T20:03:59.908261Z","iopub.execute_input":"2026-01-19T20:03:59.908639Z","iopub.status.idle":"2026-01-19T20:04:03.786183Z","shell.execute_reply.started":"2026-01-19T20:03:59.908599Z","shell.execute_reply":"2026-01-19T20:04:03.785151Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Load ECG Mask / Preprocessed Image","metadata":{}},{"cell_type":"markdown","source":"## Data\nThis notebook assumes that ECG waveform masks have already been extracted\nfrom the original images (e.g., via thresholding or segmentation).\n\nThe comparison focuses on **post-mask centerline extraction**.\n","metadata":{}},{"cell_type":"markdown","source":"## 3. Shared Preprocessing\nRed waveform extraction using HSV thresholding.\n","metadata":{}},{"cell_type":"markdown","source":"## 4. Centerline Extraction Functions\nEach version differs mainly in how the centerline is estimated and smoothed.\n","metadata":{}},{"cell_type":"code","source":"\n\ndef extract_v3(mask):\n    # cog + peak hybrid\n    H, W = mask.shape\n    y_cog = np.full(W, np.nan)\n    y_peak = np.full(W, np.nan)\n\n    for x in range(W):\n        ys = np.where(mask[:,x]>0)[0]\n        if len(ys)>0:\n            y_cog[x] = ys.mean()\n            y_peak[x] = np.median(ys)\n\n    y = np.nanmean(np.vstack([y_cog, y_peak]), axis=0)\n    return pd.Series(y).interpolate().values\n\n\ndef extract_v4(mask):\n    # v4.1 style: hybrid + median + gaussian + savgol\n    y = extract_v3(mask)\n    y = median_filter(y, size=5)\n    y = gaussian_filter1d(y, sigma=4)\n    y = savgol_filter(y, 21, 3)\n    return y\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-19T20:04:03.788617Z","iopub.execute_input":"2026-01-19T20:04:03.789086Z","iopub.status.idle":"2026-01-19T20:04:03.797837Z","shell.execute_reply.started":"2026-01-19T20:04:03.789043Z","shell.execute_reply":"2026-01-19T20:04:03.796658Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. v4 Centerlines\n\nWe compare four versions:\n\n- **v4**: Robust hybrid + denoising + smoothing\n\n","metadata":{}},{"cell_type":"markdown","source":"## 6. Normalise\n### Normalize vertical pixel positions\n### - invert y-axis to match ECG convention (upward = positive)\n### - center around mean to stabilize baseline\n### - scale to [0,1] for submission compatibility\n\n","metadata":{}},{"cell_type":"markdown","source":"## 7. Build Output datast for csv","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport cv2  # 忘れずにimport！\n\n# sample読み込み\nsample_sub = pd.read_parquet('/kaggle/input/physionet-ecg-image-digitization/sample_submission.parquet')\nsample_sub['id'] = sample_sub['id'].astype(str)\n\nprint(f\"Loaded sample_submission rows: {len(sample_sub)}\")\nprint(\"sample_sub head:\\n\", sample_sub.head())\n\ntest_df = pd.read_csv('/kaggle/input/physionet-ecg-image-digitization/test.csv')\n\n\n\nall_values = []\n\nfor _, row in test_df.iterrows():\n    record_id = row[\"id\"]\n    lead = row[\"lead\"]\n    n = int(row[\"number_of_rows\"])\n    \n    print(f\"Processing: {record_id} - {lead} (n={n})\")\n    \n    signal = np.zeros(n, dtype=np.float32)\n    \n    try:\n        img_path = f\"/kaggle/input/physionet-ecg-image-digitization/test/{record_id}.png\"\n        img = cv2.imread(img_path)\n        \n        if img is None:\n            print(f\"Image load failed: {img_path} → using zeros\")\n        else:\n            # ここからimg_hsvを定義（imgがNoneじゃない場合のみ）\n            img_hsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV)\n            \n            # 以前の範囲\n            lower_red1 = np.array([0, 30, 30])\n            upper_red1 = np.array([10, 255, 255])\n            lower_red2 = np.array([165, 30, 30])\n            upper_red2 = np.array([180, 255, 255])\n            \n                        \n            # mask作成（ORで結合）\n            mask_red = cv2.inRange(img_hsv, lower_red1, upper_red1) | \\\n                       cv2.inRange(img_hsv, lower_red2, upper_red2)\n\n  \n            \n            sig = extract_v4(mask_red)\n            \n            if sig is None or len(sig) == 0:\n                print(f\"No signal extracted → using zeros\")\n                sig = np.zeros(n)\n            \n            \n            # スケール変換（動的）\n            height = img.shape[0]\n            full_range_mv = 80.0\n            scale = full_range_mv / height\n            \n            # baselineを信号の中央値に変更（これが一番効果的）\n            if len(sig) > 0 and np.sum(~np.isnan(sig)) > 0:\n                baseline = np.nanmedian(sig)  # nanmedianでNaN無視\n            else:\n                #baseline = height / 2.0  # fallback\n                baseline = height / 2.0 + 20\n\n\n            sig = np.nan_to_num(sig, nan=baseline)\n            sig = sig[:n] if len(sig) > n else np.pad(sig, (0, n - len(sig)), mode='constant', constant_values=baseline)\n            \n            signal = (baseline - sig) * scale\n        \n\n    except Exception as e:\n        print(f\"Exception in {record_id}-{lead}: {str(e)} → fallback to zeros\")\n        signal = np.zeros(n, dtype=np.float32)\n    \n    all_values.extend(signal.tolist())\n\n\n\n# 上書き\nsample_sub['value'] = all_values\n\n# 保存\nsample_sub.to_csv('submission.csv', index=False)\n\nprint(\"Value range:\", min(all_values), \"~\", max(all_values))  # 小さいmV範囲か確認","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-19T20:18:47.537497Z","iopub.execute_input":"2026-01-19T20:18:47.537884Z","iopub.status.idle":"2026-01-19T20:18:52.243045Z","shell.execute_reply.started":"2026-01-19T20:18:47.537852Z","shell.execute_reply":"2026-01-19T20:18:52.241691Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8.sanity check","metadata":{"jupyter":{"source_hidden":true}}},{"cell_type":"markdown","source":"## Limitations\n\n- Voltage is not converted to physical units (mV)\n- Lead-wise segmentation (V1–V6) is not yet applied\n- SNR evaluation is performed separately\n\nThese steps are intentionally deferred to keep the submission pipeline simple and robust.\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}