{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":97984,"databundleVersionId":14096757}],"dockerImageVersionId":31153,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Forked from: https://www.kaggle.com/code/ambrosm/ecg-original-explained-baseline\nMany thanks for the great base notebook.\n\nThis notebook keeps the original marker detection and scanned ECG conversion pipeline, and adds a Skeleton-based Curve Tracking method for waveform extraction.\n\nIn the original baseline, the ECG waveform center is estimated by averaging the top and bottom boundaries of the black signal line. In the improved version, the black ECG waveform is skeletonized into a one-pixel-wide centerline, and this skeleton centerline is used to reconstruct the ECG signal.","metadata":{}},{"cell_type":"markdown","source":"# ECG Image Digitization with Skeleton-based Curve Tracking\n\n## Task Overview\nThe goal of this project is to convert ECG images into digital 12-lead ECG time series. This is the main task of the PhysioNet ECG Digitization competition.\n\n## Baseline Method\nThe original baseline processes high-quality scanned color ECG images (image types 3 and 11) using two steps:\n1. Detect lead endpoints using the `MarkerFinder` class.\n2. Extract ECG waveforms using the `convert_scanned_color()` function, where the waveform center is estimated by averaging the top and bottom boundaries of the black ECG line.\n\n\\[\ncenter = \\frac{top + bottom}{2}\n\\]\n\n## Proposed Innovation: Skeleton-based Curve Tracking\nThis project adds **Skeleton-based Curve Tracking** to improve waveform extraction. Instead of using `(top + bottom)/2`, skeletonization is applied to the black ECG waveform to obtain a one-pixel-wide centerline.  \nThis reduces artifacts caused by thick lines, marker occlusion, or text, and improves SNR.\n\nThe new function is:\n\n```python\nconvert_scanned_color_skeleton()","metadata":{}},{"cell_type":"markdown","source":"Image Types Used\nFocus on image types 3 and 11: high-quality scanned color images with a stable scale (~80 pixels/mV).\n\nColor separation makes it easier to distinguish black ECG curves from red gridlines.\n\nDataset Overview\nTraining set: 977 ECG records (~84 GB)\n\nEach ECG: multiple PNG images + 1 CSV ground-truth file\n\nSampling frequencies: 250, 256, 500, 512, 1000, 1025 Hz","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport cv2\nfrom glob import glob\nimport matplotlib.pyplot as plt\nfrom collections import defaultdict\nfrom tqdm import tqdm\n\nfrom skimage.morphology import skeletonize\nfrom typing import Tuple","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:11:20.295271Z","iopub.execute_input":"2026-06-02T12:11:20.295728Z","iopub.status.idle":"2026-06-02T12:11:20.300890Z","shell.execute_reply.started":"2026-06-02T12:11:20.295697Z","shell.execute_reply":"2026-06-02T12:11:20.299967Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Competition metric\n# From https://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric\nfrom typing import Tuple\n\nimport numpy as np\nimport pandas as pd\n\nimport scipy.optimize\nimport scipy.signal\n\n\nLEADS = ['I', 'II', 'III', 'aVR', 'aVL', 'aVF', 'V1', 'V2', 'V3', 'V4', 'V5', 'V6']\nMAX_TIME_SHIFT = 0.2\nPERFECT_SCORE = 384\n\n\nclass ParticipantVisibleError(Exception):\n    pass\n\n\ndef compute_power(label: np.ndarray, prediction: np.ndarray) -> Tuple[float, float]:\n    if label.ndim != 1 or prediction.ndim != 1:\n        raise ParticipantVisibleError('Inputs must be 1-dimensional arrays.')\n    finite_mask = np.isfinite(prediction)\n    if not np.any(finite_mask):\n        raise ParticipantVisibleError(\"The 'prediction' array contains no finite values (all NaN or inf).\")\n\n    prediction[~np.isfinite(prediction)] = 0\n    noise = label - prediction\n    p_signal = np.sum(label**2)\n    p_noise = np.sum(noise**2)\n    return p_signal, p_noise\n\n\ndef compute_snr(signal: float, noise: float) -> float:\n    if noise == 0:\n        # Perfect reconstruction\n        snr = PERFECT_SCORE\n    elif signal == 0:\n        snr = 0\n    else:\n        snr = min((signal / noise), PERFECT_SCORE)\n    return snr\n\n\ndef align_signals(label: np.ndarray, pred: np.ndarray, max_shift: float = float('inf')) -> np.ndarray:\n    if np.any(~np.isfinite(label)):\n        raise ParticipantVisibleError('values in label should all be finite')\n    if np.sum(np.isfinite(pred)) == 0:\n        raise ParticipantVisibleError('prediction can not all be infinite')\n\n    # Initialize the reference and digitized signals.\n    label_arr = np.asarray(label, dtype=np.float64)\n    pred_arr = np.asarray(pred, dtype=np.float64)\n\n    label_mean = np.mean(label_arr)\n    pred_mean = np.mean(pred_arr)\n\n    label_arr_centered = label_arr - label_mean\n    pred_arr_centered = pred_arr - pred_mean\n\n    # Compute the correlation between the reference and digitized signals and locate the maximum correlation.\n    correlation = scipy.signal.correlate(label_arr_centered, pred_arr_centered, mode='full')\n\n    n_label = np.size(label_arr)\n    n_pred = np.size(pred_arr)\n\n    lags = scipy.signal.correlation_lags(n_label, n_pred, mode='full')\n    valid_lags_mask = (lags >= -max_shift) & (lags <= max_shift)\n\n    max_correlation = np.nanmax(correlation[valid_lags_mask])\n    all_max_indices = np.flatnonzero(correlation == max_correlation)\n    best_idx = min(all_max_indices, key=lambda i: abs(lags[i]))\n    time_shift = lags[best_idx]\n    start_padding_len = max(time_shift, 0)\n    pred_slice_start = max(-time_shift, 0)\n    pred_slice_end = min(n_label - time_shift, n_pred)\n    end_padding_len = max(n_label - n_pred - time_shift, 0)\n    aligned_pred = np.concatenate((np.full(start_padding_len, np.nan), pred_arr[pred_slice_start:pred_slice_end], np.full(end_padding_len, np.nan)))\n\n    def objective_func(v_shift):\n        return np.nansum((label_arr - (aligned_pred - v_shift)) ** 2)\n\n    if np.any(np.isfinite(label_arr) & np.isfinite(aligned_pred)):\n        results = scipy.optimize.minimize_scalar(objective_func, method='Brent')\n        vertical_shift = results.x\n        aligned_pred -= vertical_shift\n    return aligned_pred\n\n\ndef _calculate_image_score(group: pd.DataFrame) -> float:\n    \"\"\"Helper function to calculate the total SNR score for a single image group.\"\"\"\n\n    unique_fs_values = group['fs'].unique()\n    if len(unique_fs_values) != 1:\n        raise ParticipantVisibleError('Sampling frequency should be consistent across each ecg')\n    sampling_frequency = unique_fs_values[0]\n    if sampling_frequency != int(len(group[group['lead'] == 'II']) / 10):\n        raise ParticipantVisibleError('The sequence_length should be sampling frequency * 10s')\n    sum_signal = 0\n    sum_noise = 0\n    for lead in LEADS:\n        sub = group[group['lead'] == lead]\n        label = sub['value_true'].values\n        pred = sub['value_pred'].values\n\n        aligned_pred = align_signals(label, pred, int(sampling_frequency * MAX_TIME_SHIFT))\n        p_signal, p_noise = compute_power(label, aligned_pred)\n        sum_signal += p_signal\n        sum_noise += p_noise\n    return compute_snr(sum_signal, sum_noise)\n\n\ndef score(solution: pd.DataFrame, submission: pd.DataFrame, row_id_column_name: str) -> float:\n    \"\"\"\n    Compute the mean Signal-to-Noise Ratio (SNR) across multiple ECG leads and images for the PhysioNet 2025 competition.\n    The final score is the average of the sum of SNRs over different lines, averaged over all unique images.\n    Args:\n        solution: DataFrame with ground truth values. Expected columns: 'id' and one for each lead.\n        submission: DataFrame with predicted values. Expected columns: 'id' and one for each lead.\n        row_id_column_name: The name of the unique identifier column, typically 'id'.\n    Returns:\n        The final competition score.\n\n    Examples\n    --------\n    >>> import pandas as pd\n    >>> import numpy as np\n    >>> row_id_column_name = \"id\"\n    >>> solution = pd.DataFrame({'id': ['343_0_I', '343_1_I', '343_2_I', '343_0_III', '343_1_III','343_2_III','343_0_aVR', '343_1_aVR','343_2_aVR',\\\n    '343_0_aVL', '343_1_aVL', '343_2_aVL', '343_0_aVF', '343_1_aVF','343_2_aVF','343_0_V1', '343_1_V1', '343_2_V1','343_0_V2', '343_1_V2','343_2_V2',\\\n    '343_0_V3', '343_1_V3', '343_2_V3','343_0_V4', '343_1_V4', '343_2_V4', '343_0_V5', '343_1_V5','343_2_V5','343_0_V6', '343_1_V6','343_2_V6',\\\n    '343_0_II', '343_1_II','343_2_II', '343_3_II', '343_4_II', '343_5_II','343_6_II', '343_7_II','343_8_II','343_9_II','343_10_II','343_11_II'],\\\n    'fs': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1],\\\n    'value':[0.1,0.3,0.4,0.6,0.6,0.4,0.2,0.3,0.4,0.5,0.2,0.7,0.2,0.3,0.4,0.8,0.6,0.7, 0.2,0.3,-0.1,0.5,0.6,0.7,0.2,0.9,0.4,0.5,0.6,0.7,0.1,0.3,0.4,\\\n    0.6,0.6,0.4,0.2,0.3,0.4,0.5,0.2,0.7,0.2,0.3,0.4]})\n    >>> submission = solution.copy()\n    >>> round(score(solution, submission, row_id_column_name), 4)\n    25.8433\n    >>> submission.loc[0, 'value'] = 0.9 # Introduce some noise\n    >>> round(score(solution, submission, row_id_column_name), 4)\n    13.6291\n    >>> submission.loc[4, 'value'] = 0.3 # Introduce some noise\n    >>> round(score(solution, submission, row_id_column_name), 4)\n    13.0576\n\n    >>> solution = pd.DataFrame({'id': ['343_0_I', '343_1_I', '343_2_I', '343_0_III', '343_1_III','343_2_III','343_0_aVR', '343_1_aVR','343_2_aVR',\\\n    '343_0_aVL', '343_1_aVL', '343_2_aVL', '343_0_aVF', '343_1_aVF','343_2_aVF','343_0_V1', '343_1_V1', '343_2_V1','343_0_V2', '343_1_V2','343_2_V2',\\\n    '343_0_V3', '343_1_V3', '343_2_V3','343_0_V4', '343_1_V4', '343_2_V4', '343_0_V5', '343_1_V5','343_2_V5','343_0_V6', '343_1_V6','343_2_V6',\\\n    '343_0_II', '343_1_II','343_2_II', '343_3_II', '343_4_II', '343_5_II','343_6_II', '343_7_II','343_8_II','343_9_II','343_10_II','343_11_II'],\\\n    'fs': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1],\\\n    'value':[0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]})\n    >>> round(score(solution, submission, row_id_column_name), 4)\n    -384\n    >>> submission = solution.copy()\n    >>> round(score(solution, submission, row_id_column_name), 4)\n    25.8433\n\n    >>> # test alignment\n    >>> label = np.array([0, 1, 2, 1, 0])\n    >>> pred = np.array([0, 1, 2, 1, 0])\n    >>> aligned = align_signals(label, pred)\n    >>> expected_array = np.array([0, 1, 2, 1, 0])\n    >>> np.allclose(aligned, expected_array, equal_nan=True)\n    True\n\n    >>> # Test 2: Vertical shift (DC offset) should be removed\n    >>> label = np.array([0, 1, 2, 1, 0])\n    >>> pred = np.array([10, 11, 12, 11, 10])\n    >>> aligned = align_signals(label, pred)\n    >>> expected_array = np.array([0, 1, 2, 1, 0])\n    >>> np.allclose(aligned, expected_array, equal_nan=True)\n    True\n\n    >>> # Test 3: Time shift should be corrected\n    >>> label = np.array([0, 0, 1, 2, 1, 0., 0.])\n    >>> pred = np.array([1, 2, 1, 0, 0, 0, 0])\n    >>> aligned = align_signals(label, pred)\n    >>> expected_array = np.array([np.nan, np.nan, 1, 2, 1, 0, 0])\n    >>> np.allclose(aligned, expected_array, equal_nan=True)\n    True\n    \n    >>> # Test 4: max_shift constraint prevents optimal alignment\n    >>> label = np.array([0, 0, 0, 0, 1, 2, 1]) # Peak is far\n    >>> pred = np.array([1, 2, 1, 0, 0, 0, 0])\n    >>> aligned = align_signals(label, pred, max_shift=10)\n    >>> expected_array = np.array([ np.nan, np.nan, np.nan, np.nan, 1, 2, 1])\n    >>> np.allclose(aligned, expected_array, equal_nan=True)\n    True\n\n    \"\"\"\n    for df in [solution, submission]:\n        if row_id_column_name not in df.columns:\n            raise ParticipantVisibleError(f\"'{row_id_column_name}' column not found in DataFrame.\")\n        if df['value'].isna().any():\n            raise ParticipantVisibleError('NaN exists in solution/submission')\n        if not np.isfinite(df['value']).all():\n            raise ParticipantVisibleError('Infinity exists in solution/submission')\n\n    submission = submission[['id', 'value']]\n    merged_df = pd.merge(solution, submission, on=row_id_column_name, suffixes=('_true', '_pred'))\n    merged_df['image_id'] = merged_df[row_id_column_name].str.split('_').str[0]\n    merged_df['row_id'] = merged_df[row_id_column_name].str.split('_').str[1].astype('int64')\n    merged_df['lead'] = merged_df[row_id_column_name].str.split('_').str[2]\n    merged_df.sort_values(by=['image_id', 'row_id', 'lead'], inplace=True)\n    image_scores = merged_df.groupby('image_id').apply(_calculate_image_score, include_groups=False)\n    return max(float(10 * np.log10(image_scores.mean())), -PERFECT_SCORE)","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2026-06-02T12:11:20.302284Z","iopub.execute_input":"2026-06-02T12:11:20.302536Z","iopub.status.idle":"2026-06-02T12:11:20.322787Z","shell.execute_reply.started":"2026-06-02T12:11:20.302520Z","shell.execute_reply":"2026-06-02T12:11:20.321936Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Reading Metadata and Computing the Average ECG\n\n## Metadata\nRead training and test metadata from `train.csv` and `test.csv`, which include:\n- ECG id\n- Sampling frequency\n- Lead name\n- Required output length\n\n## Average ECG Baseline\nCompute the **average waveform per lead** across the training set.  \nThis baseline is used as a fallback prediction for images that cannot be processed reliably.\n\n## Role in This Project\n- Skeleton-based extraction is applied to image types 3 and 11.\n- For unsupported images (grayscale, mobile photos), the average ECG per lead is used to ensure complete predictions.","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/physionet-ecg-image-digitization/train.csv')\ntest = pd.read_csv('/kaggle/input/physionet-ecg-image-digitization/test.csv')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:11:20.323753Z","iopub.execute_input":"2026-06-02T12:11:20.324069Z","iopub.status.idle":"2026-06-02T12:11:20.353522Z","shell.execute_reply.started":"2026-06-02T12:11:20.324054Z","shell.execute_reply":"2026-06-02T12:11:20.352967Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def fit_mean_model(train, verbose=False):\n    \"\"\"Compute minima, maxima and means of the time series\"\"\"\n    mean_dict = defaultdict(list)\n    for idx, row in tqdm(train.iterrows(), total=len(train)):\n        labels = pd.read_csv(f'/kaggle/input/physionet-ecg-image-digitization/train/{row.id}/{row.id}.csv')\n        for lead in labels.columns:\n            values = labels[lead]\n            values = values[~values.isna()]\n            mean_dict[lead].append(values)\n    \n    for lead in mean_dict.keys():\n        # Resample every time series to 20000 samples\n        mean_dict[lead] = [\n            np.interp(np.linspace(0, len(values)-1, 20000), np.arange(len(values)), values)\n            for values in mean_dict[lead]\n        ]\n\n        # Stack all ECGs\n        mean_dict[lead] = np.stack(mean_dict[lead])\n\n        # Plot the mean ECG\n        if verbose:\n            m = mean_dict[lead].mean(axis=0)\n            # s = mean_dict[lead].std(axis=0)\n            plt.figure(figsize=(12, 2))\n            plt.title(f\"Mean curve for {lead}\")\n            plt.plot(m)\n            # plt.plot(m-s/30)\n            # plt.plot(m+s/30)\n            plt.axhline(0, color='gray')\n            plt.ylabel('mV')\n            plt.show()\n\n    return mean_dict\n\ndef validate_mean_model(val, mean_dict):\n    snr_list = []\n    for idx, row in tqdm(val.iterrows(), total=len(val)):\n        labels = pd.read_csv(f'/kaggle/input/physionet-ecg-image-digitization/train/{row.id}/{row.id}.csv')\n        # Evaluate the signal-to-noise ratio\n        sum_signal = 0\n        sum_noise = 0\n        for lead in labels.columns:\n            label = labels[lead]\n            label = label[~ label.isna()]\n            pred = mean_dict[lead].mean(axis=0)\n            pred = np.interp(np.linspace(0, 1, len(label)), np.linspace(0, 1, len(pred)), pred)\n            assert len(label) == len(pred)\n    \n            aligned_pred = align_signals(label, pred, int(row.fs * MAX_TIME_SHIFT))\n            p_signal, p_noise = compute_power(label, aligned_pred)\n            sum_signal += p_signal\n            sum_noise += p_noise\n    \n        snr = compute_snr(sum_signal, sum_noise)\n        # print(f\"{idx=:4d} id={row.id} SNR: {snr:.2f}\")\n        snr_list.append(snr)\n    \n    snr = np.array(snr_list).mean()\n    val_score = max(float(10 * np.log10(snr)), -PERFECT_SCORE)\n    print(f\"# Validation SNR for mean prediction: {snr:.2f} {val_score=:.2f}\")\n\n# Validate the mean model\ntrain_test_split_loc = 780\nmean_dict = fit_mean_model(train.iloc[:train_test_split_loc], verbose=True)\nvalidate_mean_model(train.iloc[train_test_split_loc:], mean_dict)\n\n# Refit the mean model to the full dataset\nmean_dict = fit_mean_model(train, verbose=False)\n# 描述数据集，用于课程汇报\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:11:20.404436Z","iopub.execute_input":"2026-06-02T12:11:20.405389Z","iopub.status.idle":"2026-06-02T12:12:25.763208Z","shell.execute_reply.started":"2026-06-02T12:11:20.405373Z","shell.execute_reply":"2026-06-02T12:12:25.762521Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# =========================\n# Dataset Description for Presentation\n# 数据集描述：用于课程汇报\n# =========================\n\ndef describe_dataset(train, test, max_scan_samples=200):\n    print(\"Train shape:\", train.shape)\n    print(\"Test shape:\", test.shape)\n\n    print(\"\\nTrain columns:\")\n    print(train.columns.tolist())\n\n    print(\"\\nTest columns:\")\n    print(test.columns.tolist())\n\n    print(\"\\nTrain head:\")\n    display(train.head())\n\n    print(\"\\nTest head:\")\n    display(test.head())\n\n    # 训练集采样率分布\n    if 'fs' in train.columns:\n        plt.figure(figsize=(6, 4))\n        train['fs'].value_counts().sort_index().plot(kind='bar')\n        plt.title(\"Sampling Frequency Distribution in Train\")\n        plt.xlabel(\"fs\")\n        plt.ylabel(\"Count\")\n        plt.show()\n\n    # 测试集导联分布\n    if 'lead' in test.columns:\n        plt.figure(figsize=(8, 4))\n        test['lead'].value_counts().reindex(LEADS).plot(kind='bar')\n        plt.title(\"Lead Distribution in Test\")\n        plt.xlabel(\"Lead\")\n        plt.ylabel(\"Count\")\n        plt.show()\n\n    # 测试集输出长度分布\n    if 'number_of_rows' in test.columns:\n        plt.figure(figsize=(6, 4))\n        plt.hist(test['number_of_rows'], bins=30)\n        plt.title(\"Number of Rows Distribution in Test\")\n        plt.xlabel(\"number_of_rows\")\n        plt.ylabel(\"Count\")\n        plt.show()\n\n    # 统计部分训练集图片类型分布\n    image_types = []\n\n    for _, row in tqdm(train.head(max_scan_samples).iterrows(), total=min(len(train), max_scan_samples)):\n        paths = glob(f'/kaggle/input/physionet-ecg-image-digitization/train/{row.id}/{row.id}-*.png')\n        for p in paths:\n            try:\n                image_types.append(int(p[-8:-4]))\n            except:\n                pass\n\n    if len(image_types) > 0:\n        type_counts = pd.Series(image_types).value_counts().sort_index()\n\n        plt.figure(figsize=(8, 4))\n        type_counts.plot(kind='bar')\n        plt.title(f\"Image Type Distribution in First {max_scan_samples} Train ECGs\")\n        plt.xlabel(\"Image Type\")\n        plt.ylabel(\"Count\")\n        plt.show()\n\n        display(type_counts.rename(\"count\").to_frame())\n\n\ndescribe_dataset(train, test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:12:25.764759Z","iopub.execute_input":"2026-06-02T12:12:25.765045Z","iopub.status.idle":"2026-06-02T12:12:26.654066Z","shell.execute_reply.started":"2026-06-02T12:12:25.765027Z","shell.execute_reply":"2026-06-02T12:12:26.653222Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Finding Lead Endpoints\n\n## Purpose\nIdentify the start and end positions of each ECG lead before waveform extraction.\n\n## MarkerFinder\n- Detects **17 marker points** in scanned ECG images.\n- Defines boundaries for all 12 leads.\n\n## Detection Method\n- 13 markers: found using OpenCV template matching (`cv2.matchTemplate()`)\n- 4 right endpoints: estimated as linear combinations of nearby markers\n\n## Role in This Project\nBoth baseline and Skeleton-based methods use the same markers for lead segmentation. This preprocessing step remains unchanged.","metadata":{}},{"cell_type":"code","source":"class MarkerFinder:\n    \"\"\"This class finds the 13 markers in scanned ecg images and guesses the 4 line ends.\"\"\"\n    # From https://www.kaggle.com/code/ambrosm/ecg-original-explained-baseline\n    \n    def __init__(self, show_templates=False):\n        # Derive the templates from type 1 images\n        # np.max keeps the gridlines and markers and removes the ecg lines\n        ima = np.max([\n            cv2.imread('/kaggle/input/physionet-ecg-image-digitization/train/4292118763/4292118763-0001.png'),\n            cv2.imread('/kaggle/input/physionet-ecg-image-digitization/train/4289880010/4289880010-0001.png'),\n            cv2.imread('/kaggle/input/physionet-ecg-image-digitization/train/4284351157/4284351157-0001.png'),\n        ], axis=0)\n\n        # Template points in global coordinates of type 1 images\n        absolute_points = np.zeros((17, 2), dtype=int)\n        for i in range(3):\n            absolute_points[5 * i] = np.array([707 + 284 * i, 118]) # y, x\n            for j in range(1, 5):\n                absolute_points[5 * i + j] = np.array([707 + 284 * i, 118 + 492 * j])\n        absolute_points[5 * 3] = np.array([1535, 118])\n        absolute_points[5 * 3 + 1] = np.array([1535, 118 + 492 * 4])\n\n        # Top left corner of template rectangle\n        template_positions = [None] * 17\n        for i in range(len(absolute_points)):\n            if absolute_points[i][1] < 118 + 492 * 4:\n                if i % 5 == 0:\n                    template_positions[i] = (absolute_points[i][0] - 87, absolute_points[i][1] - 50) # y, x\n                else:\n                    template_positions[i] = (absolute_points[i][0] - 37, absolute_points[i][1] - 13)\n\n        # Height and width of the templates\n        template_sizes = np.array([(105, 60)] * 17) # height, width\n\n        # Convert the points to relative coordinates (inside the template)\n        template_points = [np.array([absolute_points[i][0] - template_positions[i][0],\n                                     absolute_points[i][1] - template_positions[i][1]])\n                           if template_positions[i] is not None\n                           else None\n                           for i in range(len(absolute_points))]\n\n        # Save the template matrices\n        templates = [None] * 17\n        for i in range(len(template_positions)):\n            if template_points[i] is not None:\n                template = (ima[template_positions[i][0]:template_positions[i][0]+template_sizes[i][0],\n                            template_positions[i][1]:template_positions[i][1]+template_sizes[i][1]])\n                templates[i] = template\n\n        # Plot the template matrices\n        if show_templates:\n            _, axs = plt.subplots(4, 4, figsize=(5, 7))\n            for i in range(len(template_positions)):\n                if template_points[i] is not None:\n                    template = templates[i].copy()\n                    cv2.rectangle(template,\n                                  (template_points[i][1]-1, template_points[i][0]-1),\n                                  (template_points[i][1]+1, template_points[i][0]+1), \n                                  [255, 0, 0], 2)\n                    axs[i // 5, i % 5].imshow(template)\n            for i in range(13, len(axs.ravel())):\n                axs.ravel()[i].axis('off')\n            plt.tight_layout()\n            plt.suptitle('The templates for the 13 markers', y=1.01)\n            plt.show()\n\n        self._absolute_points = absolute_points\n        self._template_positions = template_positions\n        self._template_sizes = template_sizes\n        self._template_points = template_points\n        self._templates = templates\n        \n    def find_markers(self, ima, warn=False, plot=False, title=''):\n        \"\"\"Return 17 markers as list of size-2 integer arrays (row, column)\n\n        Parameters:\n        ima: array of shape (1652, height, 3)\n        \"\"\"\n        \n        if ima.shape[0] != 1652:\n            raise ValueError(\"Implemented only for scanned images (image types 3, 4, 11, 12)\")\n\n        markers = [None] * 17\n\n        # Find 13 template-based markers\n        for j in range(len(self._templates)):\n            if self._template_points[j] is not None:\n                t = self._template_positions[j][0]-100\n                l = max(self._template_positions[j][1]-100, 0)\n                search_range = (ima[t:self._template_positions[j][0]+100+self._template_sizes[j][0],\n                                l:self._template_positions[j][1]+250+self._template_sizes[j][0]])\n                res = cv2.matchTemplate(search_range, self._templates[j], cv2.TM_CCOEFF)\n                min_val, max_val, min_loc, max_loc = cv2.minMaxLoc(res)\n    \n                top_left = max_loc\n                if warn and max_val < 3e7:\n                    bottom_right = (top_left[0] + self._templates[j].shape[1], top_left[1] + self._templates[j].shape[0])\n                    print(j, top_left, max_val)\n                    search_range = search_range.copy()\n                    cv2.rectangle(search_range, top_left, bottom_right, 0, 2)\n                    plt.imshow(search_range)\n                    plt.show()\n                markers[j] = np.array((t + top_left[1] + self._template_points[j][0], l + top_left[0] + self._template_points[j][1]))\n\n        # Guess the ends of the first three lines (can be outside the bounding box of the image)\n        for i in range(3):\n            m = markers[5 * i + 3] * 2 - markers[5 * i + 2]\n            markers[5 * i + 4] = m\n\n        # Guess the end of the fourth line (can be outside the bounding box of the image)\n        markers[16] = ((markers[14] * (284 + 260) - markers[9] * 260) / 284).astype(int)\n\n        if plot:\n            ima = ima.copy()\n            for m in markers:\n                if m is not None:\n                    cv2.rectangle(ima, (m[1]-40, m[0]-40), (m[1]+40, m[0]+40), (255, 0, 0), 2)\n            # plt.figure(figsize=(12, 8))\n            plt.imshow(ima)\n            plt.title(title)\n            plt.show()\n\n        return markers\n\n    # def baseline(self, i):\n    #     \"\"\"y coordinate of ith baseline in type 1 images\"\"\"\n    #     if i not in [0, 1, 2, 3]:\n    #         raise ValueError(\"i must be in [0, 1, 2, 3]\")\n    #     return self._absolute_points[5 * i][0]\n        \n    @staticmethod\n    def lead_info(lead):\n        \"\"\"Specify which markers mark the begin and the end of a lead.\"\"\"\n        begin, end = {\n            'I': (0, 1),\n            'II-subset': (5, 6),\n            'III': (10, 11),\n            'aVR': (1, 2),\n            'aVL': (6, 7),\n            'aVF': (11, 12),\n            'V1': (2, 3),\n            'V2': (7, 8),\n            'V3': (12, 13),\n            'V4': (3, 4),\n            'V5': (8, 9),\n            'V6': (13, 14),\n            'II': (15, 16),\n        }[lead]\n        return begin // 5, begin, end\n\n    def demo(self, ima, warn=False, title=''):\n        \"\"\"Plot the image with red markers\"\"\"\n        markers = self.find_markers(ima, warn, plot=True, title=title)\n\nmf = MarkerFinder(show_templates=False)\n\nima = cv2.imread('/kaggle/input/physionet-ecg-image-digitization/train/1026034238/1026034238-0011.png') # correct\nmf.demo(ima, warn=False, title='Scanned ECG with 17 line endpoints')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:12:26.654911Z","iopub.execute_input":"2026-06-02T12:12:26.655173Z","iopub.status.idle":"2026-06-02T12:12:27.824335Z","shell.execute_reply.started":"2026-06-02T12:12:26.655150Z","shell.execute_reply":"2026-06-02T12:12:27.823498Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Converting Scanned ECG Images\n\n## Baseline Conversion\n- Function: `convert_scanned_color()`\n- Works only for image types 3 and 11\n- Steps:\n  1. Remove most red gridlines\n  2. Sweep image from top to bottom to detect four ECG waveform lines\n  3. Use markers to segment 12 leads\n  4. Compute waveform center as `(top + bottom)/2`\n\n## Skeleton-based Conversion\n- Function: `convert_scanned_color_skeleton()`\n- Applies skeletonization to compress black ECG curves into a one-pixel-wide centerline\n- Extracts waveform along the skeleton for each lead segment\n\n## Comparison\n| Method | Centerline estimation |\n|--------|---------------------|\n| Baseline | `(top + bottom)/2` |\n| Skeleton | Skeletonized centerline |","metadata":{}},{"cell_type":"code","source":"def find_line_by_topdown_sweep(ima):\n    \"\"\"Find the topmost black line in an image and remove it.\n\n    Parameters:\n    ima: 2d boolean image array (False = black, True = white), will be updated\n\n    Return values:\n    top: topmost black pixel in every column of the matrix\n    bottom: topmost white pixel in every column of the matrix below the topmost black pixel\n    \"\"\"\n    # Find the topmost black pixels\n    top = np.argmin(ima, axis=0) # topmost black (False) pixel per column; 0 if there are no black pixels\n\n    # Paint black everything above\n    mask = np.tile(np.arange(len(ima)).reshape(-1, 1), reps=(1, ima.shape[1]))\n    mask = mask >= top # True for lower part of image\n    ima &= mask # paint black whatever is above the line\n\n    # Find the topmost white pixels\n    bottom = np.argmax(ima, axis=0) # topmost white (True) pixel per column; 0 if there are no white pixels\n\n    # Paint white everything above\n    bottomx = np.maximum(bottom, np.median(top) + 100) # overpaints the letters\n    mask = np.tile(np.arange(len(ima)).reshape(-1, 1), reps=(1, ima.shape[1]))\n    mask = mask < bottomx # True for upper part of image\n    ima |= mask # paint white whatever is above the line\n    ima[:,:-1] |= mask[:,1:]\n    ima[:,1:] |= mask[:,:-1]\n\n    return top, bottom\n\ndef get_lead_from_top_bottom(tops, bottoms, lead, number_of_rows, markers):\n    \"\"\"Extract and resample one lead from an ECG line.\n    \n    Parameters:\n    tops: list of 4 arrays of shape (image_width, )\n    bottoms: list of 4 arrays of shape (image_width, )\n    lead: one of the 12 lead labels (string)\n    number_of_rows: number of samples required (int)\n    markers: 17 markers as list of size-2 integer arrays (row, column)\n    \"\"\"\n\n    # Select the markers and determine the baseline\n    line, begin, end = mf.lead_info(lead)\n    top = tops[line]\n    bottom = bottoms[line]\n    begin, end = markers[begin], markers[end]\n    baseline = np.linspace(begin[0], end[0], end[1] - begin[1])\n\n    pred0 = (top[begin[1]:end[1]] + bottom[begin[1]:end[1]]) / 2\n    if len(pred0) < len(baseline):\n        print(f\"smaller: {len(pred0)} < {len(baseline)}\")\n    baseline = baseline[:len(pred0)] # in case end is outside the image\n    pred = baseline - pred0\n\n    # Scale\n    pred /= 80 # 80 pixels = 1 mV\n\n    # Fix pixels obscured by the markers\n    if lead in ['aVR', 'aVL', 'aVF', 'V1', 'V2', 'V3', 'V4', 'V5', 'V6']:\n        # first four values can be obscured by the marker\n        pred[:4][pred[:4] > 0.2] = pred[4]\n    if lead in ['I', 'II-subset', 'III', 'aVR', 'aVL', 'aVF', 'V1', 'V2', 'V3']:\n        # last five pixels can be obscured by the marker\n        pred[-5:][pred[-5:] > 0.2] = pred[-6]\n    if lead in ['I', 'II-subset', 'III', 'II']:\n        # first pixel can be obscured by the marker\n        if 0.9 < pred[0] and pred[0] < 1.1 and pred[1] < 0.5:\n            pred[0] = pred[1]\n\n    # Upsample\n    pred = np.interp(np.linspace(0, 1, number_of_rows),\n                     np.linspace(0, 1, len(pred)),\n                     pred)\n    \n    # Fix implausible predictions\n    pred = np.where(np.abs(pred) <= 0.9, pred, 0)\n        \n    return pred\n\ndef convert_scanned_color(ima, markers, n_timesteps, verbose=False):\n    \"\"\"Convert a scanned color image (type 3 or 11) to 12 leads.\n\n    The function first extracts the four lines from the image. As\n    the four lines have nonnegligible width, we construct two lists:\n    - tops = y coordinates of the topmost black pixels in the lines\n    - bottoms = y coordinates of the topmost white pixels below the lines\n    Either list is a list of 4 arrays of shape (image_width, )\n\n    Parameters:\n    ima: 3-channel BGR image with height 1652 and width ≈2200.\n    markers: 17 markers as list of size-2 integer arrays (row, column)\n    n_timesteps: number of samples required per lead (dict)\n\n    Returns:\n    preds: dict with 12 time series\n    \"\"\"\n    # Crop the image and convert to black and white\n    # We use only the red channel (channel 2) so that the red gridlines disappear\n    # False = black, True = white\n    # The text at the top of the image is discarded.\n    crop_top = 400\n    ima = ima[crop_top:, :, 2] > 160\n\n    # Denoise single and double black pixels\n    iima = ima.astype(np.uint8)\n    ima = (iima[:-2, :-2] + iima[:-2, 1:-1] + iima[:-2, 2:]\n           + iima[1:-1, :-2] + iima[1:-1, 1:-1] + iima[1:-1, 2:]\n           + iima[2:, :-2] + iima[2:, 1:-1] + iima[2:, 2:]) >= 7\n\n    # Plot the denoised black-and-white image\n    if verbose:\n        plt.figure(figsize=(6, 4))\n        plt.imshow(ima)\n        plt.title('Denoised black-and-white')\n        # plt.savefig('ima.png')\n        plt.show()\n    \n    # Find the four lines\n    tops, bottoms = [], []\n    for i in range(4):\n        top, bottom = find_line_by_topdown_sweep(ima)\n        tops.append(top)\n        bottoms.append(bottom)\n\n    # Transform to global coordinates\n    tops = [t + crop_top for t in tops]\n    bottoms = [b + crop_top for b in bottoms]\n\n    # Extract the twelve leads from the four lines\n    # (as the first part of II is duplicated, we extract it twice\n    # and take the average)\n    n_timesteps['II-subset'] = n_timesteps['I']\n    preds = {}\n    for i, lead in enumerate(LEADS + ['II-subset']):\n        pred = get_lead_from_top_bottom(tops, bottoms, lead, n_timesteps[lead], markers)\n        preds[lead] = pred\n\n    preds['II'][:len(preds['II-subset'])] = (preds['II'][:len(preds['II-subset'])] + preds['II-subset']) / 2\n    del preds['II-subset']\n\n    # === ★ ===\n    for lead, sig in preds.items():\n        sig = np.asarray(sig)\n        n = len(sig)\n        if lead not in n_timesteps:\n            continue\n        replace_len = int(n_timesteps[lead] / 125)\n\n        if replace_len > 0 and replace_len * 2 < n:\n            # 始端部分を置換\n            sig[:replace_len] = sig[replace_len]\n            # 終端部分を置換\n            sig[-replace_len:] = sig[-replace_len - 1]\n\n        preds[lead] = sig\n    # === ★ ===\n\n    # Apply Einthoven's law\n    apply_einthoven(preds)\n\n    return preds\n\ndef apply_einthoven(preds):\n    \"\"\"Apply Einthoven's law to improve the predictions.\n    \n    Parameters:\n    pred: dict of time series, will be updated\n    \"\"\"\n    residual = preds['I'] + preds['III'] - preds['II'][:len(preds['III'])]\n    correction = residual / 3\n    preds['I'] -= correction\n    preds['III'] -= correction\n    preds['II'][:len(preds['III'])] += correction\n    \n    residual = preds['aVR'] + preds['aVL'] + preds['aVF']\n    correction = residual / 3\n    preds['aVR'] -= correction\n    preds['aVL'] -= correction\n    preds['aVF'] -= correction\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:12:27.825798Z","iopub.execute_input":"2026-06-02T12:12:27.826086Z","iopub.status.idle":"2026-06-02T12:12:27.844379Z","shell.execute_reply.started":"2026-06-02T12:12:27.826067Z","shell.execute_reply":"2026-06-02T12:12:27.843774Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ==========================================================\n# Skeleton-based Curve Tracking\n# 新增创新点：用骨架化提取 ECG 曲线中心线\n# ==========================================================\n\nSKELETON_CONFIG = {\n    \"crop_top\": 400,\n    \"pixel_per_mv\": 80,\n    \"edge_fix_ratio\": 125,\n    \"skeleton_band_margin\": 80,\n    \"skeleton_min_points\": 3,\n}\n\n\ndef extract_skeleton_curve_from_roi(roi_bw, baseline_local=None, min_points=3):\n    \"\"\"\n    从局部 ROI 中提取 ECG 黑色曲线的 skeleton 中心线。\n\n    roi_bw:\n        True = white, False = black\n\n    baseline_local:\n        ROI 内部的基线位置。\n        如果同一列出现多个骨架点，就选择离基线最近的点。\n    \"\"\"\n\n    # skeletonize 需要前景为 True\n    # 原图里黑色 ECG 线是 False，所以这里取反\n    foreground = ~roi_bw\n    skel = skeletonize(foreground)\n\n    h, w = skel.shape\n    curve = np.full(w, np.nan, dtype=float)\n\n    for x in range(w):\n        ys = np.where(skel[:, x])[0]\n\n        if len(ys) >= min_points:\n            if baseline_local is not None and x < len(baseline_local):\n                curve[x] = ys[np.argmin(np.abs(ys - baseline_local[x]))]\n            else:\n                curve[x] = np.mean(ys)\n\n        elif len(ys) > 0:\n            curve[x] = np.mean(ys)\n\n    # 没有骨架点的列用线性插值补齐\n    nans = np.isnan(curve)\n\n    if np.all(nans):\n        curve[:] = h / 2\n\n    elif np.any(nans):\n        valid = ~nans\n        curve[nans] = np.interp(\n            np.where(nans)[0],\n            np.where(valid)[0],\n            curve[valid]\n        )\n\n    return curve, skel\n\n\ndef get_lead_from_skeleton(\n    bw_clean,\n    lead,\n    number_of_rows,\n    markers,\n    crop_top=400,\n    pixel_per_mv=80,\n    band_margin=80,\n    min_points=3,\n    return_debug=False,\n):\n    \"\"\"\n    使用 Skeleton-based Curve Tracking 提取单个导联。\n    \"\"\"\n\n    line, begin_idx, end_idx = mf.lead_info(lead)\n    begin, end = markers[begin_idx], markers[end_idx]\n\n    img_h, img_w = bw_clean.shape\n\n    x0 = int(max(0, begin[1]))\n    x1 = int(min(img_w, end[1]))\n\n    if x1 <= x0 + 5:\n        pred = np.zeros(number_of_rows)\n        if return_debug:\n            return pred, None\n        return pred\n\n    xs = np.arange(x0, x1)\n\n    baseline_global = begin[0] + (end[0] - begin[0]) * (xs - begin[1]) / max((end[1] - begin[1]), 1)\n\n    y0_global = int(np.floor(np.min(baseline_global) - band_margin))\n    y1_global = int(np.ceil(np.max(baseline_global) + band_margin))\n\n    y0 = max(0, y0_global - crop_top)\n    y1 = min(img_h, y1_global - crop_top)\n\n    if y1 <= y0 + 5:\n        pred = np.zeros(number_of_rows)\n        if return_debug:\n            return pred, None\n        return pred\n\n    roi_bw = bw_clean[y0:y1, x0:x1]\n    baseline_local = baseline_global - crop_top - y0\n\n    curve_local, skel = extract_skeleton_curve_from_roi(\n        roi_bw,\n        baseline_local=baseline_local,\n        min_points=min_points\n    )\n\n    curve_global = curve_local + y0 + crop_top\n\n    baseline_global = baseline_global[:len(curve_global)]\n    pred = baseline_global - curve_global\n\n    # 像素转 mV，和原始 baseline 保持一致\n    pred = pred / pixel_per_mv\n\n    # 修复 marker 遮挡造成的边界异常\n    if lead in ['aVR', 'aVL', 'aVF', 'V1', 'V2', 'V3', 'V4', 'V5', 'V6']:\n        if len(pred) > 5:\n            pred[:4][pred[:4] > 0.2] = pred[4]\n\n    if lead in ['I', 'II-subset', 'III', 'aVR', 'aVL', 'aVF', 'V1', 'V2', 'V3']:\n        if len(pred) > 6:\n            pred[-5:][pred[-5:] > 0.2] = pred[-6]\n\n    if lead in ['I', 'II-subset', 'III', 'II']:\n        if len(pred) > 1:\n            if 0.9 < pred[0] < 1.1 and pred[1] < 0.5:\n                pred[0] = pred[1]\n\n    # 重采样到目标长度\n    pred = np.interp(\n        np.linspace(0, 1, number_of_rows),\n        np.linspace(0, 1, len(pred)),\n        pred\n    )\n\n    # 去除明显不合理值\n    pred = np.where(np.abs(pred) <= 0.9, pred, 0)\n\n    if return_debug:\n        debug = {\n            \"roi_bw\": roi_bw,\n            \"skeleton\": skel,\n            \"curve_local\": curve_local,\n            \"baseline_local\": baseline_local,\n            \"lead\": lead,\n        }\n        return pred, debug\n\n    return pred\n\n\ndef convert_scanned_color_skeleton(\n    ima,\n    markers,\n    n_timesteps,\n    verbose=False,\n    skeleton_config=None,\n    return_debug=False,\n):\n    \"\"\"\n    Skeleton 版本的 scanned color ECG 转换函数。\n\n    原始方法：\n        pred0 = (top + bottom) / 2\n\n    新方法：\n        对 ECG 黑色曲线做 skeletonize，\n        提取单像素中心线作为曲线位置。\n    \"\"\"\n\n    if skeleton_config is None:\n        skeleton_config = SKELETON_CONFIG\n\n    crop_top = skeleton_config[\"crop_top\"]\n\n    # 和原始方法一致：裁掉顶部文字，只使用红色通道去掉红色网格\n    bw = ima[crop_top:, :, 2] > 160\n\n    # 和原始方法一致：3x3 去噪\n    iima = bw.astype(np.uint8)\n    bw_clean = (\n        iima[:-2, :-2] + iima[:-2, 1:-1] + iima[:-2, 2:]\n        + iima[1:-1, :-2] + iima[1:-1, 1:-1] + iima[1:-1, 2:]\n        + iima[2:, :-2] + iima[2:, 1:-1] + iima[2:, 2:]\n    ) >= 7\n\n    if verbose:\n        plt.figure(figsize=(6, 4))\n        plt.imshow(bw_clean, cmap='gray')\n        plt.title('Skeleton method: denoised binary image')\n        plt.show()\n\n    # 保留 top-down sweep，用于流程一致和调试\n    # 注意 find_line_by_topdown_sweep 会修改输入图，所以这里用 copy\n    bw_for_sweep = bw_clean.copy()\n    tops, bottoms = [], []\n\n    for i in range(4):\n        top, bottom = find_line_by_topdown_sweep(bw_for_sweep)\n        tops.append(top + crop_top)\n        bottoms.append(bottom + crop_top)\n\n    n_timesteps = n_timesteps.copy()\n    n_timesteps['II-subset'] = n_timesteps['I']\n\n    preds = {}\n    debug_info = {}\n\n    for lead in LEADS + ['II-subset']:\n        need_debug = return_debug and lead in ['I', 'II', 'V1', 'V5']\n\n        result = get_lead_from_skeleton(\n            bw_clean,\n            lead,\n            n_timesteps[lead],\n            markers,\n            crop_top=crop_top,\n            pixel_per_mv=skeleton_config[\"pixel_per_mv\"],\n            band_margin=skeleton_config[\"skeleton_band_margin\"],\n            min_points=skeleton_config[\"skeleton_min_points\"],\n            return_debug=need_debug,\n        )\n\n        if need_debug:\n            pred, dbg = result\n            debug_info[lead] = dbg\n        else:\n            pred = result\n\n        preds[lead] = pred\n\n    # II 长导联和 II-subset 融合\n    preds['II'][:len(preds['II-subset'])] = (\n        preds['II'][:len(preds['II-subset'])] + preds['II-subset']\n    ) / 2\n    del preds['II-subset']\n\n    # Edge Fix，和原始方法保持一致\n    for lead, sig in preds.items():\n        sig = np.asarray(sig)\n        n = len(sig)\n\n        if lead not in n_timesteps:\n            continue\n\n        replace_len = int(n_timesteps[lead] / skeleton_config[\"edge_fix_ratio\"])\n\n        if replace_len > 0 and replace_len * 2 < n:\n            sig[:replace_len] = sig[replace_len]\n            sig[-replace_len:] = sig[-replace_len - 1]\n\n        preds[lead] = sig\n\n    # Einthoven 定律修正，和原始方法保持一致\n    apply_einthoven(preds)\n\n    if return_debug:\n        return preds, debug_info\n\n    return preds\n\n\ndef visualize_skeleton_debug(debug_info):\n    \"\"\"\n    可视化 Skeleton 提取过程。\n    这部分图可以直接放进课程汇报的定性实验。\n    \"\"\"\n\n    for lead, dbg in debug_info.items():\n        if dbg is None:\n            continue\n\n        roi_bw = dbg[\"roi_bw\"]\n        skel = dbg[\"skeleton\"]\n        curve = dbg[\"curve_local\"]\n        baseline = dbg[\"baseline_local\"]\n\n        plt.figure(figsize=(12, 3))\n\n        plt.subplot(1, 3, 1)\n        plt.imshow(roi_bw, cmap='gray')\n        plt.title(f'{lead}: ROI binary')\n        plt.axis('off')\n\n        plt.subplot(1, 3, 2)\n        plt.imshow(skel, cmap='gray')\n        plt.title(f'{lead}: Skeleton')\n        plt.axis('off')\n\n        plt.subplot(1, 3, 3)\n        plt.imshow(roi_bw, cmap='gray')\n        plt.plot(curve, linewidth=1, label='Skeleton curve')\n        plt.plot(baseline, linewidth=1, label='Baseline')\n        plt.title(f'{lead}: Curve tracking')\n        plt.legend()\n        plt.axis('off')\n\n        plt.tight_layout()\n        plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:12:27.845168Z","iopub.execute_input":"2026-06-02T12:12:27.845406Z","iopub.status.idle":"2026-06-02T12:12:27.871954Z","shell.execute_reply.started":"2026-06-02T12:12:27.845380Z","shell.execute_reply":"2026-06-02T12:12:27.871390Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# =========================\n# Skeleton Demo\n# 用于生成课程汇报中的定性可视化图\n# =========================\n\ndemo_row = train.iloc[100]\n\ndemo_paths = sorted(\n    glob(f'/kaggle/input/physionet-ecg-image-digitization/train/{demo_row.id}/{demo_row.id}-*.png')\n)\n\ndemo_path = None\nfor p in demo_paths:\n    img_type = int(p[-8:-4])\n    if img_type in [3, 11]:\n        demo_path = p\n        break\n\nif demo_path is not None:\n    demo_ima = cv2.imread(demo_path)\n\n    demo_markers = mf.find_markers(\n        demo_ima,\n        plot=True,\n        title='Demo image with 17 markers'\n    )\n\n    demo_labels = pd.read_csv(\n        f'/kaggle/input/physionet-ecg-image-digitization/train/{demo_row.id}/{demo_row.id}.csv'\n    )\n\n    demo_n_timesteps = {\n        lead: (~demo_labels[lead].isna()).sum()\n        for lead in LEADS\n    }\n\n    demo_preds, demo_debug = convert_scanned_color_skeleton(\n        demo_ima,\n        demo_markers,\n        demo_n_timesteps,\n        verbose=True,\n        return_debug=True\n    )\n\n    visualize_skeleton_debug(demo_debug)\nelse:\n    print(\"No type 3 or type 11 image found for demo.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:12:27.872883Z","iopub.execute_input":"2026-06-02T12:12:27.873173Z","iopub.status.idle":"2026-06-02T12:12:30.028181Z","shell.execute_reply.started":"2026-06-02T12:12:27.873155Z","shell.execute_reply":"2026-06-02T12:12:30.027552Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Validation\n\n## Purpose\nEvaluate waveform extraction methods locally before generating final submission.\n\n## Procedure\nFor selected training images:\n1. Detect 17 markers\n2. Convert image to 12 ECG signals (baseline vs Skeleton)\n3. Compare with ground-truth CSV signals\n4. Compute Signal-to-Noise Ratio (SNR)\n5. Plot `y_true` vs `y_pred` for qualitative assessment\n\n## Experiments\n- Baseline: `convert_scanned_color()`\n- Skeleton: `convert_scanned_color_skeleton()`\n\n## Output\n- Quantitative: SNR values\n- Qualitative: waveform plots\n- Supports: ablation study and failure case analysis","metadata":{}},{"cell_type":"code","source":"def validate_algorithm(train, image_types, convert):\n    \"\"\"Convert a few training images, plot the output and compute the signal-to-noise ratio\"\"\"\n    snr_list = []\n    index_list = []\n    is_first_ecg = True\n    for idx, row in train.iterrows():\n        # print(idx, row.id)\n        labels = pd.read_csv(f'/kaggle/input/physionet-ecg-image-digitization/train/{row.id}/{row.id}.csv')\n        png_paths = sorted(glob(f'/kaggle/input/physionet-ecg-image-digitization/train/{row.id}/{row.id}-*.png'))\n        for path in png_paths:\n            img_type = int(path[-8:-4])\n            # ima = cv2.imread(path)\n            # shape = ima.shape\n            # assert (img_type == 1) <= (shape == (1700, 2200, 3)) # 200 pixels per inch on Letter paper\n            # assert (img_type == 3) <= (shape[0] == 1652)\n            # assert (img_type == 4) <= (shape[0] == 1652)\n            # assert (img_type == 5) <= (len(shape) == 3 and shape[2] == 3)\n            # assert (img_type == 6) <= ((shape == (4000, 3000, 3) or (shape == (3000, 4000, 3))))\n            # assert (img_type == 9) <= ((shape == (3024, 4032, 3) or (shape == (4032, 3024, 3))))\n            # assert (img_type == 10) <= ((shape == (3024, 4032, 3) or (shape == (4032, 3024, 3))))\n            # assert (img_type == 11) <= (shape[0] == 1652)\n            # assert (img_type == 12) <= (shape[0] == 1652)\n            if img_type in image_types:\n                ima = cv2.imread(path)\n\n                # Find the 17 line endpoints\n                markers = mf.find_markers(ima, plot=is_first_ecg, title='Image with 17 markers')\n\n                # Convert the image to 12 leads\n                n_timesteps = {lead: (~ labels[lead].isna()).sum() for lead in LEADS}\n                preds = convert(ima, markers, n_timesteps, verbose=is_first_ecg)\n                \n                # Evaluate the signal-to-noise ratio, plot y_true vs. y_pred\n                if is_first_ecg:\n                    _, axs = plt.subplots(6, 2, figsize=(12, 18))\n                sum_signal = 0\n                sum_noise = 0\n                for i, lead in enumerate(LEADS):\n                    label = labels[lead]\n                    label = label[~ label.isna()]\n                    pred = preds[lead]\n            \n                    aligned_pred = align_signals(label, pred, int(row.fs * MAX_TIME_SHIFT))\n                    p_signal, p_noise = compute_power(label, aligned_pred)\n                    sum_signal += p_signal\n                    sum_noise += p_noise\n    \n                    if is_first_ecg:\n                        ax = axs.T.ravel()[i]\n                        ax.set_title(lead)\n                        ax.plot(label.values, label='y_true')\n                        ax.plot(pred, label='y_pred')\n                        ax.set_xlabel('timestep')\n                        ax.set_ylabel('mV')\n                        ax.legend()\n                if is_first_ecg:\n                    plt.tight_layout()\n                    plt.suptitle('y_true vs. y_pred', y=1.01)\n                    plt.show()\n                snr = compute_snr(sum_signal, sum_noise)\n                print(f\"{idx=:4d} id={row.id:10d} {img_type=:2d} SNR: {snr:5.2f}\")\n                snr_list.append(snr)\n                index_list.append([idx, img_type])\n    \n                if is_first_ecg:\n                    print('\\n')\n            else:\n                snr_list.append(1)\n                index_list.append([idx, img_type])\n        is_first_ecg = False\n    \n    snr = np.array(snr_list).mean()\n    val_score = max(float(10 * np.log10(snr)), -PERFECT_SCORE)\n    print(f\"# Average SNR: {snr:.2f} {val_score=:.2f}\")\n    snr_df = pd.DataFrame(index_list, columns=['idx', 'type'])\n    snr_df['snr'] = snr_list\n    snr_df.to_csv('~snr.csv', index=False)\n\n# Baseline 定量 + 定性实验\nvalidate_algorithm(train.iloc[100:110], image_types=[3, 11], convert=convert_scanned_color)\n\n# Skeleton 定量 + 定性实验\nvalidate_algorithm(train.iloc[100:110], image_types=[3, 11], convert=convert_scanned_color_skeleton)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:12:30.029026Z","iopub.execute_input":"2026-06-02T12:12:30.029596Z","iopub.status.idle":"2026-06-02T12:13:01.342105Z","shell.execute_reply.started":"2026-06-02T12:12:30.029577Z","shell.execute_reply":"2026-06-02T12:13:01.341314Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# =========================\n# Ablation Study\n# 消融实验：Baseline vs Skeleton\n# =========================\n\ndef run_ablation_experiment(train_subset, image_types=[3, 11]):\n    results = []\n\n    methods = [\n        (\"Baseline\", convert_scanned_color),\n        (\"Skeleton\", convert_scanned_color_skeleton),\n    ]\n\n    for method_name, convert_func in methods:\n        print(f\"\\n========== Running {method_name} ==========\")\n\n        snr_list = []\n\n        for idx, row in tqdm(train_subset.iterrows(), total=len(train_subset)):\n            labels = pd.read_csv(\n                f'/kaggle/input/physionet-ecg-image-digitization/train/{row.id}/{row.id}.csv'\n            )\n\n            png_paths = sorted(\n                glob(f'/kaggle/input/physionet-ecg-image-digitization/train/{row.id}/{row.id}-*.png')\n            )\n\n            for path in png_paths:\n                img_type = int(path[-8:-4])\n\n                if img_type not in image_types:\n                    continue\n\n                try:\n                    ima = cv2.imread(path)\n                    markers = mf.find_markers(ima)\n\n                    n_timesteps = {\n                        lead: (~labels[lead].isna()).sum()\n                        for lead in LEADS\n                    }\n\n                    preds = convert_func(\n                        ima,\n                        markers,\n                        n_timesteps,\n                        verbose=False\n                    )\n\n                    sum_signal = 0\n                    sum_noise = 0\n\n                    for lead in LEADS:\n                        label = labels[lead]\n                        label = label[~label.isna()]\n                        pred = preds[lead]\n\n                        aligned_pred = align_signals(\n                            label.values,\n                            pred,\n                            int(row.fs * MAX_TIME_SHIFT)\n                        )\n\n                        p_signal, p_noise = compute_power(\n                            label.values,\n                            aligned_pred\n                        )\n\n                        sum_signal += p_signal\n                        sum_noise += p_noise\n\n                    snr = compute_snr(sum_signal, sum_noise)\n                    snr_list.append(snr)\n\n                except Exception as e:\n                    print(f\"[FAILED] {method_name}, idx={idx}, path={path}, error={e}\")\n\n        avg_snr = np.mean(snr_list)\n\n        results.append({\n            \"Method\": method_name,\n            \"Average SNR\": avg_snr,\n            \"Validation Score\": max(float(10 * np.log10(avg_snr)), -PERFECT_SCORE)\n        })\n\n    result_df = pd.DataFrame(results)\n    display(result_df)\n\n    plt.figure(figsize=(6, 4))\n    plt.bar(result_df[\"Method\"], result_df[\"Average SNR\"])\n    plt.title(\"Ablation Study: Baseline vs Skeleton\")\n    plt.xlabel(\"Method\")\n    plt.ylabel(\"Average SNR\")\n    plt.show()\n\n    result_df.to_csv(\"ablation_baseline_vs_skeleton.csv\", index=False)\n\n    return result_df\n\n\nablation_df = run_ablation_experiment(train.iloc[100:110], image_types=[3, 11])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:13:01.343073Z","iopub.execute_input":"2026-06-02T12:13:01.343511Z","iopub.status.idle":"2026-06-02T12:13:20.913035Z","shell.execute_reply.started":"2026-06-02T12:13:01.343490Z","shell.execute_reply":"2026-06-02T12:13:20.912118Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# =========================\n# Parameter Sensitivity Analysis\n# 参数敏感性分析：skeleton_band_margin\n# =========================\n\ndef run_sensitivity_experiment(train_subset, image_types=[3, 11], margins=[40, 60, 80, 100, 120]):\n    records = []\n\n    for margin in margins:\n        print(f\"\\n========== skeleton_band_margin = {margin} ==========\")\n\n        config = SKELETON_CONFIG.copy()\n        config[\"skeleton_band_margin\"] = margin\n\n        def convert_with_margin(ima, markers, n_timesteps, verbose=False):\n            return convert_scanned_color_skeleton(\n                ima,\n                markers,\n                n_timesteps,\n                verbose=verbose,\n                skeleton_config=config,\n                return_debug=False\n            )\n\n        snr_list = []\n\n        for idx, row in tqdm(train_subset.iterrows(), total=len(train_subset)):\n            labels = pd.read_csv(\n                f'/kaggle/input/physionet-ecg-image-digitization/train/{row.id}/{row.id}.csv'\n            )\n\n            png_paths = sorted(\n                glob(f'/kaggle/input/physionet-ecg-image-digitization/train/{row.id}/{row.id}-*.png')\n            )\n\n            for path in png_paths:\n                img_type = int(path[-8:-4])\n\n                if img_type not in image_types:\n                    continue\n\n                try:\n                    ima = cv2.imread(path)\n                    markers = mf.find_markers(ima)\n\n                    n_timesteps = {\n                        lead: (~labels[lead].isna()).sum()\n                        for lead in LEADS\n                    }\n\n                    preds = convert_with_margin(\n                        ima,\n                        markers,\n                        n_timesteps,\n                        verbose=False\n                    )\n\n                    sum_signal = 0\n                    sum_noise = 0\n\n                    for lead in LEADS:\n                        label = labels[lead]\n                        label = label[~label.isna()]\n                        pred = preds[lead]\n\n                        aligned_pred = align_signals(\n                            label.values,\n                            pred,\n                            int(row.fs * MAX_TIME_SHIFT)\n                        )\n\n                        p_signal, p_noise = compute_power(\n                            label.values,\n                            aligned_pred\n                        )\n\n                        sum_signal += p_signal\n                        sum_noise += p_noise\n\n                    snr = compute_snr(sum_signal, sum_noise)\n                    snr_list.append(snr)\n\n                except Exception as e:\n                    print(f\"[FAILED] margin={margin}, idx={idx}, path={path}, error={e}\")\n\n        avg_snr = np.mean(snr_list)\n\n        records.append({\n            \"skeleton_band_margin\": margin,\n            \"Average SNR\": avg_snr,\n            \"Validation Score\": max(float(10 * np.log10(avg_snr)), -PERFECT_SCORE)\n        })\n\n    sensitivity_df = pd.DataFrame(records)\n    display(sensitivity_df)\n\n    plt.figure(figsize=(6, 4))\n    plt.plot(\n        sensitivity_df[\"skeleton_band_margin\"],\n        sensitivity_df[\"Average SNR\"],\n        marker='o'\n    )\n    plt.title(\"Sensitivity Analysis: skeleton_band_margin\")\n    plt.xlabel(\"skeleton_band_margin\")\n    plt.ylabel(\"Average SNR\")\n    plt.grid(True)\n    plt.show()\n\n    sensitivity_df.to_csv(\"sensitivity_skeleton_band_margin.csv\", index=False)\n\n    return sensitivity_df\n\n\nsensitivity_df = run_sensitivity_experiment(\n    train.iloc[100:110],\n    image_types=[3, 11],\n    margins=[40, 60, 80, 100, 120]\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:13:20.914025Z","iopub.execute_input":"2026-06-02T12:13:20.914679Z","iopub.status.idle":"2026-06-02T12:14:14.124279Z","shell.execute_reply.started":"2026-06-02T12:13:20.914654Z","shell.execute_reply":"2026-06-02T12:14:14.123665Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# =========================\n# Failure Case Analysis\n# 失败案例展示：找 Skeleton 方法 SNR 最低的样本\n# =========================\n\ndef collect_skeleton_snr_cases(train_subset, image_types=[3, 11]):\n    records = []\n\n    for idx, row in tqdm(train_subset.iterrows(), total=len(train_subset)):\n        labels = pd.read_csv(\n            f'/kaggle/input/physionet-ecg-image-digitization/train/{row.id}/{row.id}.csv'\n        )\n\n        png_paths = sorted(\n            glob(f'/kaggle/input/physionet-ecg-image-digitization/train/{row.id}/{row.id}-*.png')\n        )\n\n        for path in png_paths:\n            img_type = int(path[-8:-4])\n\n            if img_type not in image_types:\n                continue\n\n            try:\n                ima = cv2.imread(path)\n                markers = mf.find_markers(ima)\n\n                n_timesteps = {\n                    lead: (~labels[lead].isna()).sum()\n                    for lead in LEADS\n                }\n\n                preds = convert_scanned_color_skeleton(\n                    ima,\n                    markers,\n                    n_timesteps,\n                    verbose=False\n                )\n\n                sum_signal = 0\n                sum_noise = 0\n\n                for lead in LEADS:\n                    label = labels[lead]\n                    label = label[~label.isna()]\n                    pred = preds[lead]\n\n                    aligned_pred = align_signals(\n                        label.values,\n                        pred,\n                        int(row.fs * MAX_TIME_SHIFT)\n                    )\n\n                    p_signal, p_noise = compute_power(\n                        label.values,\n                        aligned_pred\n                    )\n\n                    sum_signal += p_signal\n                    sum_noise += p_noise\n\n                snr = compute_snr(sum_signal, sum_noise)\n\n                records.append({\n                    \"idx\": idx,\n                    \"id\": row.id,\n                    \"type\": img_type,\n                    \"path\": path,\n                    \"snr\": snr\n                })\n\n            except Exception as e:\n                records.append({\n                    \"idx\": idx,\n                    \"id\": row.id,\n                    \"type\": img_type,\n                    \"path\": path,\n                    \"snr\": 0,\n                    \"error\": str(e)\n                })\n\n    return pd.DataFrame(records)\n\n\ndef visualize_failure_case(case_row):\n    idx = int(case_row[\"idx\"])\n    row = train.loc[idx]\n\n    labels = pd.read_csv(\n        f'/kaggle/input/physionet-ecg-image-digitization/train/{row.id}/{row.id}.csv'\n    )\n\n    ima = cv2.imread(case_row[\"path\"])\n\n    markers = mf.find_markers(\n        ima,\n        plot=True,\n        title=f\"Failure Case Markers | id={row.id} | SNR={case_row['snr']:.2f}\"\n    )\n\n    n_timesteps = {\n        lead: (~labels[lead].isna()).sum()\n        for lead in LEADS\n    }\n\n    preds, debug_info = convert_scanned_color_skeleton(\n        ima,\n        markers,\n        n_timesteps,\n        verbose=True,\n        return_debug=True\n    )\n\n    visualize_skeleton_debug(debug_info)\n\n    _, axs = plt.subplots(6, 2, figsize=(12, 18))\n\n    for i, lead in enumerate(LEADS):\n        label = labels[lead]\n        label = label[~label.isna()]\n        pred = preds[lead]\n\n        ax = axs.T.ravel()[i]\n        ax.set_title(lead)\n        ax.plot(label.values, label='y_true')\n        ax.plot(pred, label='y_pred')\n        ax.set_xlabel('timestep')\n        ax.set_ylabel('mV')\n        ax.legend()\n\n    plt.tight_layout()\n    plt.suptitle(\n        f\"Failure Case: id={row.id}, type={int(case_row['type'])}, SNR={case_row['snr']:.2f}\",\n        y=1.01\n    )\n    plt.show()\n\n\nfailure_df = collect_skeleton_snr_cases(\n    train.iloc[100:110],\n    image_types=[3, 11]\n)\n\nworst_cases = failure_df.sort_values(\"snr\").head(3)\ndisplay(worst_cases)\n\nfor _, case in worst_cases.iterrows():\n    visualize_failure_case(case)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:14:14.126393Z","iopub.execute_input":"2026-06-02T12:14:14.126829Z","iopub.status.idle":"2026-06-02T12:14:37.328407Z","shell.execute_reply.started":"2026-06-02T12:14:14.126811Z","shell.execute_reply":"2026-06-02T12:14:37.327451Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Converting Test Images for Submission\n\n## Test-time Strategy\n- For image types 3 and 11: use waveform extraction\n  - Baseline: `convert_scanned_color()`\n  - Skeleton: `convert_scanned_color_skeleton()` (controlled by `USE_SKELETON_FOR_SUBMISSION`)\n- For unsupported images: use average waveform per lead\n\n## Final Output\n- Predicted values for all test images are saved to `submission.csv`\n- Guarantees complete coverage while improving accuracy for high-quality scans\n\n## Benefits of Skeleton-based Submission\n- More precise waveform extraction\n- Reduced artifacts near markers or edges\n- Smoother ECG curves\n- Improved SNR compared to baseline extraction","metadata":{}},{"cell_type":"code","source":"def is_color_image(ima):\n    \"\"\" Test if a 3-channel image has colors.\"\"\"\n    return ima.std(axis=2).mean() != 0","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:14:37.329300Z","iopub.execute_input":"2026-06-02T12:14:37.329585Z","iopub.status.idle":"2026-06-02T12:14:37.333493Z","shell.execute_reply.started":"2026-06-02T12:14:37.329568Z","shell.execute_reply":"2026-06-02T12:14:37.332674Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"USE_SKELETON_FOR_SUBMISSION = True\n\nsubmission_data = []\nold_id = None\npreds = None\n\nfor idx, row in test.iterrows():\n    if row.id != old_id:\n        path = f\"/kaggle/input/physionet-ecg-image-digitization/test/{row.id}.png\"\n\n        ima = cv2.imread(path)\n\n        if ima is not None:\n            shape = ima.shape\n            good_shape = shape[0] == 1652  # scanned images have 1652 rows\n\n            if good_shape and is_color_image(ima):\n                # Find the 17 line endpoints\n                markers = mf.find_markers(ima)\n\n                # Convert the image to 12 time series\n                n_timesteps = {\n                    lead: row.fs * 10 if lead == 'II' else row.fs * 10 // 4\n                    for lead in LEADS\n                }\n\n                if USE_SKELETON_FOR_SUBMISSION:\n                    preds = convert_scanned_color_skeleton(\n                        ima,\n                        markers,\n                        n_timesteps,\n                        verbose=False\n                    )\n                else:\n                    preds = convert_scanned_color(\n                        ima,\n                        markers,\n                        n_timesteps,\n                        verbose=False\n                    )\n\n            else:\n                # we cannot interpret the image -> predict the mean\n                preds = None\n\n        else:\n            # image reading failed\n            preds = None\n\n        old_id = row.id\n\n    if row.lead == 'II':\n        assert row.number_of_rows == row.fs * 10\n    else:\n        assert row.number_of_rows == row.fs * 10 // 4\n\n    if preds is not None:\n        pred = preds[row.lead]\n    else:\n        pred = mean_dict[row.lead].mean(axis=0)\n        pred = np.interp(\n            np.linspace(0, 1, row.number_of_rows),\n            np.linspace(0, 1, len(pred)),\n            pred\n        )\n\n    assert len(pred) == row.number_of_rows\n\n    for timestep in range(row.number_of_rows):\n        signal_id = f\"{row.id}_{timestep}_{row.lead}\"\n\n        submission_data.append({\n            'id': signal_id,\n            'value': pred[timestep]\n        })\n\nsubmission_df = pd.DataFrame(submission_data)\n\nprint(f\"Length: {len(submission_df)}\")\n\nsubmission_df.to_csv('submission.csv', index=False)\n\n!head submission.csv","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-02T12:14:37.334480Z","iopub.execute_input":"2026-06-02T12:14:37.335031Z","iopub.status.idle":"2026-06-02T12:14:38.888454Z","shell.execute_reply.started":"2026-06-02T12:14:37.335014Z","shell.execute_reply":"2026-06-02T12:14:38.887757Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}