{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":97984,"databundleVersionId":14096757,"sourceType":"competition"}],"dockerImageVersionId":31153,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Takeaways\n\n1) Some of the images of the printouts are not aligned with ordinary images axis-wise. For example, some signals are placed vertically in the image such that they should be read from top to bottom.  \n2) Point *1)* is extended to the fact that there are more than two cases of how signals are placed in the image, regarding its aspect ratio. For example, 1. signals from left to right & image width larger than height, 2. signals from left to right & image height larger than width, 3. signals from top to bottom & image height larger than width(in `'/kaggle/input/physionet-ecg-image-digitization/train/2042290760/2042290760-0006.png'`). The 4th case could be found, but I haven't yet noticed it. These can be found, especially, in mobile-captured images.  \n3) Some of the signals in `0001` images extend their coverage down to the text `25mm/s` and `10mm/mV`, and up to the top of the image. Thus, cropping the upper and lower parts might not be a reasonable approach for pre-processing. In particular, `V6` lead seems to have very outstanding outlier in its values, which can be found in `Section 7)`  ","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom pathlib import Path\nfrom PIL import Image\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\n\nimport os \nimport random\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\n# import os\n# for dirname, _, filenames in os.walk('/kaggle/input'):\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:08:56.536825Z","iopub.execute_input":"2025-10-29T03:08:56.537116Z","iopub.status.idle":"2025-10-29T03:08:57.683624Z","shell.execute_reply.started":"2025-10-29T03:08:56.537093Z","shell.execute_reply":"2025-10-29T03:08:57.682776Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1) Set the random seed","metadata":{}},{"cell_type":"code","source":"# Reference: https://www.kaggle.com/code/rhythmcam/random-seed-everything?scriptVersionId=77985442&cellId=2\n# basic random seed\nDEFAULT_RANDOM_SEED = 1105\n\ndef seedBasic(seed=DEFAULT_RANDOM_SEED):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    \n# tensorflow random seed \nimport tensorflow as tf \ndef seedTF(seed=DEFAULT_RANDOM_SEED):\n    tf.random.set_seed(seed)\n    \n# torch random seed\nimport torch\ndef seedTorch(seed=DEFAULT_RANDOM_SEED):\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n      \n# basic + tensorflow + torch \ndef seedEverything(seed=DEFAULT_RANDOM_SEED):\n    seedBasic(seed)\n    seedTF(seed)\n    seedTorch(seed)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:08:57.684802Z","iopub.execute_input":"2025-10-29T03:08:57.685163Z","iopub.status.idle":"2025-10-29T03:09:05.856170Z","shell.execute_reply.started":"2025-10-29T03:08:57.685140Z","shell.execute_reply":"2025-10-29T03:09:05.855420Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2) Set the paths","metadata":{}},{"cell_type":"code","source":"root_str = \"/kaggle/input/physionet-ecg-image-digitization\"\nroot = Path(root_str)\nfolder_train = root / \"train\"\nfolder_test = root / \"test\"\npath_train_csv = root / \"train.csv\"\npath_test_csv = root / \"test.csv\"\n\ndf_train = pd.read_csv(path_train_csv)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:09:05.857000Z","iopub.execute_input":"2025-10-29T03:09:05.857526Z","iopub.status.idle":"2025-10-29T03:09:05.868935Z","shell.execute_reply.started":"2025-10-29T03:09:05.857497Z","shell.execute_reply":"2025-10-29T03:09:05.867996Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3) Load data","metadata":{}},{"cell_type":"code","source":"def load_file_paths(root_folder: Path, suffix: str = None):\n    file_paths = root_folder.glob(\"**/*\")\n    result = []\n    for p in file_paths:\n\n        # files only\n        if p.is_file():\n\n            # if condition not met, skip\n            if suffix is not None and not p.name.endswith(suffix):\n                continue\n            result.append(p)\n    return result                ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:09:05.871354Z","iopub.execute_input":"2025-10-29T03:09:05.871782Z","iopub.status.idle":"2025-10-29T03:09:05.884750Z","shell.execute_reply.started":"2025-10-29T03:09:05.871752Z","shell.execute_reply":"2025-10-29T03:09:05.883944Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_images(\n    paths: list[Path | str],\n    num_row: int,\n    num_col: int,\n    fig_size: tuple[int, int] = None\n):\n    fig, axes = plt.subplots(\n        nrows=num_row, \n        ncols=num_col,\n        # sharex=\"all\",\n        # sharey=\"all\",\n    )\n    if fig_size is None:\n        factor = 3.5\n        fig_size = (factor * num_col, factor * num_row)\n    fig.set_size_inches(fig_size)\n\n    for i in range(num_row):\n        for j in range(num_col):\n            img_idx = i*num_row + j\n            axes[i][j].imshow(Image.open(paths[img_idx]))\n    \n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:09:05.885595Z","iopub.execute_input":"2025-10-29T03:09:05.886615Z","iopub.status.idle":"2025-10-29T03:09:05.900921Z","shell.execute_reply.started":"2025-10-29T03:09:05.886585Z","shell.execute_reply":"2025-10-29T03:09:05.899910Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"paths_0001 = load_file_paths(folder_train, \"0001.png\")\npaths_0003 = load_file_paths(folder_train, \"0003.png\")\npaths_0004 = load_file_paths(folder_train, \"0004.png\")\npaths_0005 = load_file_paths(folder_train, \"0005.png\")\npaths_0006 = load_file_paths(folder_train, \"0006.png\")\npaths_0009 = load_file_paths(folder_train, \"0009.png\")\npaths_0010 = load_file_paths(folder_train, \"0010.png\")\npaths_0011 = load_file_paths(folder_train, \"0011.png\")\npaths_0012 = load_file_paths(folder_train, \"0012.png\")\npaths_csv = load_file_paths(folder_train, \".csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:09:05.901926Z","iopub.execute_input":"2025-10-29T03:09:05.902187Z","iopub.status.idle":"2025-10-29T03:09:33.213348Z","shell.execute_reply.started":"2025-10-29T03:09:05.902151Z","shell.execute_reply":"2025-10-29T03:09:33.212638Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4) Analysis","metadata":{}},{"cell_type":"markdown","source":"#### Shape consistency","metadata":{}},{"cell_type":"code","source":"# Data preparation\nshapes = {}\naspect_ratios = {}\npaths_group = [\n    paths_0001, paths_0003, paths_0004, paths_0005, paths_0006,\n    paths_0009, paths_0010, paths_0011, paths_0012\n]\n\nfor idx, paths in enumerate(paths_group):\n    i = 0\n    tmp_type = \"none\"\n    for p in paths:\n        if i == 0:\n            tmp_type = p.name[-8:-4]\n            shapes[tmp_type] = list()\n            aspect_ratios[tmp_type] = list()\n            i += 1\n            \n        tmp_img = Image.open(p)\n        tmp_ar = tmp_img.height / tmp_img.width\n        shapes[tmp_type].append((tmp_img.height, tmp_img.width))\n        aspect_ratios[tmp_type].append(tmp_ar)\n    print(f\"{idx}\", end=\" \")\nprint(\"Getting shape information done!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:09:33.214109Z","iopub.execute_input":"2025-10-29T03:09:33.214404Z","iopub.status.idle":"2025-10-29T03:09:44.139633Z","shell.execute_reply.started":"2025-10-29T03:09:33.214383Z","shell.execute_reply":"2025-10-29T03:09:44.138751Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# set the variables\nnunique_shapes = pd.DataFrame(shapes).nunique()\nnunique_ars = pd.DataFrame(aspect_ratios).nunique()\nars = np.array([vs for vs in aspect_ratios.values()]).reshape(-1)\nmin_val = min(nunique_shapes.min(), nunique_ars.min()) - 1\nmax_val = max(nunique_shapes.max(), nunique_ars.max()) + 1\n\n# plot\nfig, axes = plt.subplots(2, 1)\nfig.set_size_inches(10, 8)\n\n# -- 1) Overview\nnunique_shapes.plot(kind=\"bar\", color=\"orange\", ax=axes[0], label=\"Shape\")\nnunique_ars.plot(kind=\"bar\", color=\"skyblue\", ax=axes[0], alpha=0.5, label=\"AR\")\naxes[0].set_xlabel(\"Image Type\")\naxes[0].set_ylabel(\"Unique Count\")\naxes[0].set_yticks(range(min_val, max_val + 1))\naxes[0].legend()\naxes[0].axhline(1, color=\"red\", linestyle=\"--\", alpha=0.5)\naxes[0].text(len(nunique_shapes), 1, \"Above this line has multiple kinds within its type\", color=\"red\", alpha=0.7)\naxes[0].set_title(\"Unique Count Per Image Type\")\n\n# -- 2) Aspect Ratio distribution\naxes[1].hist(ars, bins=500)\naxes[1].set_xlabel(\"AR values\")\naxes[1].set_xticks(\n    ticks=np.linspace(min(ars), max(ars), 20).round(2), \n    labels=np.linspace(min(ars), max(ars), 20).round(2), \n    rotation=45\n)\naxes[1].set_ylabel(\"AR Count\")\naxes[1].set_title(\"AR Distribution\")\naxes[1].legend()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:09:44.140530Z","iopub.execute_input":"2025-10-29T03:09:44.140860Z","iopub.status.idle":"2025-10-29T03:09:45.497223Z","shell.execute_reply.started":"2025-10-29T03:09:44.140831Z","shell.execute_reply":"2025-10-29T03:09:45.496372Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Images with AR `1.33` are found quite a lot. This means there are rotated images","metadata":{}},{"cell_type":"markdown","source":"#### Image examples","metadata":{}},{"cell_type":"code","source":"num_to_show = 8\nselected_paths = np.asarray(paths_0001)\nselected_indices = np.random.choice(np.arange(len(selected_paths)), num_to_show)\nplot_images(\n    selected_paths[selected_indices], \n    2, \n    num_to_show // 2 if num_to_show % 2 == 0 else num_to_show // 2 + 1\n)\n\n# Print the size\nrep_idx = selected_indices[0]\nrep_img = Image.open(selected_paths[rep_idx])\nprint()\nprint(f\"Selected indices: {selected_indices}\")\nprint(f\"Shape and Aspect Ratio for the image with index [{rep_idx}]\")\nprint(f\"Shape: {rep_img.size}\", f\"Aspect Ratio: {rep_img.height / rep_img.width:.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:09:45.498317Z","iopub.execute_input":"2025-10-29T03:09:45.498584Z","iopub.status.idle":"2025-10-29T03:09:49.701176Z","shell.execute_reply.started":"2025-10-29T03:09:45.498564Z","shell.execute_reply":"2025-10-29T03:09:49.700244Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_to_show = 8\nselected_paths = np.asarray(paths_0003)\nselected_indices = np.random.choice(np.arange(len(selected_paths)), num_to_show)\nplot_images(\n    selected_paths[selected_indices], \n    2, \n    num_to_show // 2 if num_to_show % 2 == 0 else num_to_show // 2 + 1\n)\n\n# Print the size\nrep_idx = selected_indices[0]\nrep_img = Image.open(selected_paths[rep_idx])\nprint()\nprint(f\"Selected indices: {selected_indices}\")\nprint(f\"Shape and Aspect Ratio for the image with index [{rep_idx}]\")\nprint(f\"Shape: {rep_img.size}\", f\"Aspect Ratio: {rep_img.height / rep_img.width:.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:09:49.704091Z","iopub.execute_input":"2025-10-29T03:09:49.704421Z","iopub.status.idle":"2025-10-29T03:09:55.054496Z","shell.execute_reply.started":"2025-10-29T03:09:49.704397Z","shell.execute_reply":"2025-10-29T03:09:55.053559Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_to_show = 8\nselected_paths = np.asarray(paths_0004)\nselected_indices = np.random.choice(np.arange(len(selected_paths)), num_to_show)\nplot_images(\n    selected_paths[selected_indices], \n    2, \n    num_to_show // 2 if num_to_show % 2 == 0 else num_to_show // 2 + 1\n)\n\n# Print the size\nrep_idx = selected_indices[0]\nrep_img = Image.open(selected_paths[rep_idx])\nprint()\nprint(f\"Selected indices: {selected_indices}\")\nprint(f\"Shape and Aspect Ratio for the image with index [{rep_idx}]\")\nprint(f\"Shape: {rep_img.size}\", f\"Aspect Ratio: {rep_img.height / rep_img.width:.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:09:55.055518Z","iopub.execute_input":"2025-10-29T03:09:55.055780Z","iopub.status.idle":"2025-10-29T03:09:59.717328Z","shell.execute_reply.started":"2025-10-29T03:09:55.055760Z","shell.execute_reply":"2025-10-29T03:09:59.716448Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_to_show = 8\nselected_paths = np.asarray(paths_0005)\nselected_indices = np.random.choice(np.arange(len(selected_paths)), num_to_show)\nplot_images(\n    selected_paths[selected_indices], \n    2, \n    num_to_show // 2 if num_to_show % 2 == 0 else num_to_show // 2 + 1\n)\n\n# Print the size\nrep_idx = selected_indices[0]\nrep_img = Image.open(selected_paths[rep_idx])\nprint()\nprint(f\"Selected indices: {selected_indices}\")\nprint(f\"Shape and Aspect Ratio for the image with index [{rep_idx}]\")\nprint(f\"Shape: {rep_img.size}\", f\"Aspect Ratio: {rep_img.height / rep_img.width:.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:09:59.718426Z","iopub.execute_input":"2025-10-29T03:09:59.719090Z","iopub.status.idle":"2025-10-29T03:10:15.260210Z","shell.execute_reply.started":"2025-10-29T03:09:59.719063Z","shell.execute_reply":"2025-10-29T03:10:15.259382Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_to_show = 8\nselected_paths = np.asarray(paths_0006)\nselected_indices = np.random.choice(np.arange(len(selected_paths)), num_to_show)\nplot_images(\n    selected_paths[selected_indices], \n    2, \n    num_to_show // 2 if num_to_show % 2 == 0 else num_to_show // 2 + 1\n)\n\n# Print the size\nrep_idx = selected_indices[0]\nrep_img = Image.open(selected_paths[rep_idx])\nprint()\nprint(f\"Selected indices: {selected_indices}\")\nprint(f\"Shape and Aspect Ratio for the image with index [{rep_idx}]\")\nprint(f\"Shape: {rep_img.size}\", f\"Aspect Ratio: {rep_img.height / rep_img.width:.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:10:15.261119Z","iopub.execute_input":"2025-10-29T03:10:15.261450Z","iopub.status.idle":"2025-10-29T03:10:31.209856Z","shell.execute_reply.started":"2025-10-29T03:10:15.261429Z","shell.execute_reply":"2025-10-29T03:10:31.208816Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_to_show = 8\nselected_paths = np.asarray(paths_0009)\nselected_indices = np.random.choice(np.arange(len(selected_paths)), num_to_show)\nplot_images(\n    selected_paths[selected_indices], \n    2, \n    num_to_show // 2 if num_to_show % 2 == 0 else num_to_show // 2 + 1\n)\n\n# Print the size\nrep_idx = selected_indices[0]\nrep_img = Image.open(selected_paths[rep_idx])\nprint()\nprint(f\"Selected indices: {selected_indices}\")\nprint(f\"Shape and Aspect Ratio for the image with index [{rep_idx}]\")\nprint(f\"Shape: {rep_img.size}\", f\"Aspect Ratio: {rep_img.height / rep_img.width:.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:10:31.210954Z","iopub.execute_input":"2025-10-29T03:10:31.211216Z","iopub.status.idle":"2025-10-29T03:10:46.862835Z","shell.execute_reply.started":"2025-10-29T03:10:31.211196Z","shell.execute_reply":"2025-10-29T03:10:46.861885Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_to_show = 8\nselected_paths = np.asarray(paths_0010)\nselected_indices = np.random.choice(np.arange(len(selected_paths)), num_to_show)\nplot_images(\n    selected_paths[selected_indices], \n    2, \n    num_to_show // 2 if num_to_show % 2 == 0 else num_to_show // 2 + 1\n)\n\n# Print the size\nrep_idx = selected_indices[0]\nrep_img = Image.open(selected_paths[rep_idx])\nprint()\nprint(f\"Selected indices: {selected_indices}\")\nprint(f\"Shape and Aspect Ratio for the image with index [{rep_idx}]\")\nprint(f\"Shape: {rep_img.size}\", f\"Aspect Ratio: {rep_img.height / rep_img.width:.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:10:46.863832Z","iopub.execute_input":"2025-10-29T03:10:46.864445Z","iopub.status.idle":"2025-10-29T03:11:01.365250Z","shell.execute_reply.started":"2025-10-29T03:10:46.864422Z","shell.execute_reply":"2025-10-29T03:11:01.364316Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_to_show = 8\nselected_paths = np.asarray(paths_0011)\nselected_indices = np.random.choice(np.arange(len(selected_paths)), num_to_show)\nplot_images(\n    selected_paths[selected_indices], \n    2, \n    num_to_show // 2 if num_to_show % 2 == 0 else num_to_show // 2 + 1\n)\n\n# Print the size\nrep_idx = selected_indices[0]\nrep_img = Image.open(selected_paths[rep_idx])\nprint()\nprint(f\"Selected indices: {selected_indices}\")\nprint(f\"Shape and Aspect Ratio for the image with index [{rep_idx}]\")\nprint(f\"Shape: {rep_img.size}\", f\"Aspect Ratio: {rep_img.height / rep_img.width:.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:11:01.366497Z","iopub.execute_input":"2025-10-29T03:11:01.366770Z","iopub.status.idle":"2025-10-29T03:11:06.848888Z","shell.execute_reply.started":"2025-10-29T03:11:01.366749Z","shell.execute_reply":"2025-10-29T03:11:06.847924Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_to_show = 8\nselected_paths = np.asarray(paths_0012)\nselected_indices = np.random.choice(np.arange(len(selected_paths)), num_to_show)\nplot_images(\n    selected_paths[selected_indices], \n    2, \n    num_to_show // 2 if num_to_show % 2 == 0 else num_to_show // 2 + 1\n)\n\n# Print the size\nrep_idx = selected_indices[0]\nrep_img = Image.open(selected_paths[rep_idx])\nprint()\nprint(f\"Selected indices: {selected_indices}\")\nprint(f\"Shape and Aspect Ratio for the image with index [{rep_idx}]\")\nprint(f\"Shape: {rep_img.size}\", f\"Aspect Ratio: {rep_img.height / rep_img.width:.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:11:06.849867Z","iopub.execute_input":"2025-10-29T03:11:06.850136Z","iopub.status.idle":"2025-10-29T03:11:11.450167Z","shell.execute_reply.started":"2025-10-29T03:11:06.850116Z","shell.execute_reply":"2025-10-29T03:11:11.449008Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5) Signal Foorprints","metadata":{}},{"cell_type":"code","source":"# set the example\nrep_grey = Image.open(paths_0001[0]).convert(\"L\")\nheatmap = np.zeros_like(rep_grey, dtype=bool)\nthreshold_from_bg = 75\n\n# gather information\nfor p in paths_0001:\n    tmp_img_arr = np.array(Image.open(p).convert(\"L\")).astype(float)\n    tmp_img_arr = np.where(tmp_img_arr < threshold_from_bg, True, False).astype(bool)\n    heatmap += tmp_img_arr\n\n# display the example\nt = np.array(rep_grey)\nt = np.where(t < threshold_from_bg, t, 255)\nplt.figure()\nplt.title(\"Example of How To Extract Signals From BG\")\nplt.imshow(t)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:11:11.451262Z","iopub.execute_input":"2025-10-29T03:11:11.451573Z","iopub.status.idle":"2025-10-29T03:12:37.238304Z","shell.execute_reply.started":"2025-10-29T03:11:11.451552Z","shell.execute_reply":"2025-10-29T03:12:37.237341Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"heatmap_reverse = np.abs(1 - heatmap)\nplt.figure(figsize=(20, 10))\nplt.title(\"All Footprints of Signals From 0001 Images\")\nsns.heatmap(heatmap_reverse)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:12:37.239546Z","iopub.execute_input":"2025-10-29T03:12:37.239999Z","iopub.status.idle":"2025-10-29T03:12:42.401538Z","shell.execute_reply.started":"2025-10-29T03:12:37.239969Z","shell.execute_reply":"2025-10-29T03:12:42.400733Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6) How To Construct `0001` From the Raw Signals","metadata":{}},{"cell_type":"code","source":"# set the example id\nexample_id = \"1063816858\"\nexample_dir = folder_train / example_id\nexample_img = example_dir / (example_id + \"-0001.png\")\nexample_csv = example_dir / (example_id + \".csv\")\n\ndf_example = pd.read_csv(example_csv)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:12:42.402678Z","iopub.execute_input":"2025-10-29T03:12:42.403302Z","iopub.status.idle":"2025-10-29T03:12:42.420903Z","shell.execute_reply.started":"2025-10-29T03:12:42.403250Z","shell.execute_reply":"2025-10-29T03:12:42.419926Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# gather data\nlead_order = [\n    [\"I\", \"aVR\", \"V1\", \"V4\"],\n    [\"II\", \"aVL\", \"V2\", \"V5\"],\n    [\"III\", \"aVF\", \"V3\", \"V6\"],\n    [\"II\"]\n]\n\nexample = [[] for _ in range(4)]\nexample_fs = len(df_example) / 10\nhalf_sig_len = int(example_fs * 5)\nis_fs_odd = example_fs % 2 == 1\n\nfor row_idx, row in enumerate(lead_order):\n    for lead in row:\n        mask = df_example[lead].isna()\n        values = df_example[lead][~mask]\n        if lead == \"II\" and row_idx == 1:\n            values = values[:half_sig_len // 2]\n                \n        example[row_idx].extend(values.to_list())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:12:42.421834Z","iopub.execute_input":"2025-10-29T03:12:42.422084Z","iopub.status.idle":"2025-10-29T03:12:42.435601Z","shell.execute_reply.started":"2025-10-29T03:12:42.422064Z","shell.execute_reply":"2025-10-29T03:12:42.434697Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# plot them\nfig, axes = plt.subplot_mosaic(\"04;14;24;34\")\nfig.set_size_inches(16, 6)\ny_lim_max = 2\ny_lim_min = -2\nfor i in range(5):\n    if i < 4:\n        axes[str(i)].plot(example[i])\n        axes[str(i)].set_ylim(y_lim_min, y_lim_max)\n        if i == 0:\n            axes[str(i)].set_title(\"Signals From Raw Data\")\n    else:\n        axes[str(i)].imshow(Image.open(example_img))\n        axes[str(i)].set_title(\"Signals From Original Image\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:12:42.436520Z","iopub.execute_input":"2025-10-29T03:12:42.436861Z","iopub.status.idle":"2025-10-29T03:12:43.666074Z","shell.execute_reply.started":"2025-10-29T03:12:42.436839Z","shell.execute_reply":"2025-10-29T03:12:43.664873Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7) Min, Max Values For Each Lead","metadata":{}},{"cell_type":"code","source":"ending_leads = [row[-1] for row in lead_order]\nmin_dict = {}\nmax_dict = {}\nfor p in paths_csv:\n    df_tmp = pd.read_csv(p)\n    for row_idx, row in enumerate(lead_order):\n\n        # skip the redundant check\n        if row_idx == 3:\n            continue\n\n        for lead in row:\n            if min_dict.get(lead) is None:\n                min_dict[lead] = []\n            if max_dict.get(lead) is None:\n                max_dict[lead] = []\n            min_dict[lead].append(df_tmp[lead].min())\n            max_dict[lead].append(df_tmp[lead].max())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:12:43.667216Z","iopub.execute_input":"2025-10-29T03:12:43.667613Z","iopub.status.idle":"2025-10-29T03:12:50.311680Z","shell.execute_reply.started":"2025-10-29T03:12:43.667590Z","shell.execute_reply":"2025-10-29T03:12:50.310667Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Min and Max Values for Each Lead\\n\")\nfor row_idx, row in enumerate(lead_order):\n    if row_idx == 3:\n        continue\n    for lead in row:\n        print(f\"{lead.rjust(4)} = Min : {min(min_dict[lead])} / Max : {max(max_dict[lead])}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-29T03:12:50.312762Z","iopub.execute_input":"2025-10-29T03:12:50.313035Z","iopub.status.idle":"2025-10-29T03:12:50.319660Z","shell.execute_reply.started":"2025-10-29T03:12:50.313014Z","shell.execute_reply":"2025-10-29T03:12:50.318658Z"}},"outputs":[],"execution_count":null}]}