{
  "id": 662842,
  "title": "How to digitize ECG waveforms after neural network–based extraction",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/662842",
  "author_name": "",
  "post_date": "2025-12-15T06:35:14.217930200Z",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am trying to digitize ECG waveforms from images after extracting the waveform region using a neural network (segmentation model).</p>\n<p>While I can obtain a binary mask or a cleaned waveform image from the model, I am not sure about the correct or standard way to convert this image-based waveform into a 1D digital signal (time series).</p>\n<p>I would really appreciate it if anyone could share a recommended pipeline or best practices for digitizing waveforms from ECG images after NN-based segmentation.</p>\n<p>Thank you very much!</p>",
  "messages": [
    {
      "id": "3376801",
      "postDate": "12/15/2025 06:35:14",
      "content": "<p>I am trying to digitize ECG waveforms from images after extracting the waveform region using a neural network (segmentation model).</p>\n<p>While I can obtain a binary mask or a cleaned waveform image from the model, I am not sure about the correct or standard way to convert this image-based waveform into a 1D digital signal (time series).</p>\n<p>I would really appreciate it if anyone could share a recommended pipeline or best practices for digitizing waveforms from ECG images after NN-based segmentation.</p>\n<p>Thank you very much!</p>",
      "rawMarkdown": "I am trying to digitize ECG waveforms from images after extracting the waveform region using a neural network (segmentation model).\n\nWhile I can obtain a binary mask or a cleaned waveform image from the model, I am not sure about the correct or standard way to convert this image-based waveform into a 1D digital signal (time series).\n\nI would really appreciate it if anyone could share a recommended pipeline or best practices for digitizing waveforms from ECG images after NN-based segmentation.\n\nThank you very much!",
      "votes": null
    },
    {
      "id": "3378328",
      "postDate": "12/17/2025 22:33:17",
      "content": "<p>I wouldn't say this is best (or even good) practice or a standard pipeline, however, something that can work is the Viterbi's algorithm, although this is not an algorithm you'd normally use here, it more likely has applications elsewhere (NLP and error correction in telecom signals)</p>\n<p>The idea is this:</p>\n<p>you have a binary mask of size (H x W), with 0s corresponding to background and 1s corresponding to foreground (your signal segment). For each column along the second axis (axis=1), find the centroid (or farthest point from baseline) of contiguous 1s, you will use these centroids to construct a DAG (Directed Acyclic Graph), with the centroids being the nodes and connection between centroids of consecutive steps being the edges, but you will not store all of the graph nodes and edges, rather you do this:</p>\n<p>1: For each centroid at position t, associate it with all centroids at t-1 and compute a cost function (I will explain a 'good enough' cost function later). EG: suppose position t has 4 contiguous 1s, that is 4 centroids, if position t-1 had 3 centroids, now you will match each centroid at t to all centroids at t-1, forming 12 edges or paths.</p>\n<p>2: For each of the centroids in t, get the centroid from t-1 that best optimizes the cost function, in other words, get the edge with the least cost for a given centroid. That will be the centroid from the previous timestep that leads to the current timestep, you store the centroids at timestep t, as well as their corresponding best previous centroids from each timestep and the corresponding cost values. Do note that the cost values from the previous timesteps are propagated into the future by adding them to the currently computed cost, EG: cost_t = cost_{t-1} + D(y_t, y_{t-1}), where D is our base cost.</p>\n<p>3: continue this process until you get to the very last timestep t=T-1</p>\n<p>4: At t=T-1, you backtrack. You do this by getting the centroid with the least cost value at t=T-1, let's call that y_{T-1}, then you retrieve the best centroid from T-2 that lead to y_{T-1}, which will be y_{T-2}, then you retrieve the best centroid from T-3 that lead to y_{T-2}, which is y_{T-3}, you do this until you get to t=0. And that is the Viterbi's solution to this problem. It selects the path of least resistance (or least cost if you may), This is somewhat equivalent to constructing the full graph and running an A-star algorithm to find the best path, but it is far more efficient.</p>\n<p>This is an example code that still needs a lot of improvement:</p>\n<pre><code>def interpolate_nan(array: np.ndarray, inplace: bool=False) -&gt; np.ndarray:\n    if not inplace:\n        array = array.copy()\n    nans = np.isnan(array)\n    indices = np.arange(len(array))\n    xp = indices[not_nans := ~nans] \n    fp = array[not_nans]\n    array[nans] = np.interp(indices[nans], xp, fp)\n    return array\n\ndef viterbi(\n        mask: np.ndarray, \n        row_mm_center: float,\n        row_mm_spacing: float,\n        hpx_spacing: float,\n        vpx_spacing: float,\n        vmm_spacing: float=5,\n        hcrop_offset: int=8,\n        mm_per_mv: float=10,\n        alpha: float=0.5,\n        mm_tol: float=0.0,\n        node_type: Literal[\"centroid\", \"farthest\"]=\"centroid\"\n    ):\n    assert node_type in [\"centroid\", \"farthest\"]\n    assert alpha &gt;= 0 and alpha &lt;= 1 \n\n    mm2px_along_y = lambda mm : mm * (vpx_spacing / vmm_spacing)\n    px2mm_along_y = lambda px : px * (vmm_spacing / vpx_spacing)\n    mm2mv = lambda mm : mm / mm_per_mv\n    dist_cost = lambda x1, x2: np.abs(x1 - x2)\n    smoothness_cost = lambda xt, xtm1, xtm2: np.abs((xt - xtm1) - (xtm1 - xtm2))\n\n    y_center = mm2px_along_y(row_mm_center)\n    row_spacing = mm2px_along_y(row_mm_spacing)\n    tol = mm2px_along_y(mm_tol)\n    threshold = row_spacing / 2 + tol\n\n    _, xs = np.where(mask)\n    mask = mask[:, xs.min()+int(hpx_spacing)+hcrop_offset : xs.max()]\n    W = mask.shape[1]\n\n    all_curr_states = []\n    all_prev_indexes = []\n    all_costs = []\n\n    for x in range(0, W):\n        m = mask[:, x]\n        if np.where(m)[0].shape[0] == 0:\n            all_curr_states.append(np.asarray([]))\n            all_prev_indexes.append(np.asarray([]))\n            all_costs.append(np.asarray([]))\n            continue\n\n        cc = cc3d.connected_components(m[:, None])\n        cc_stats = cc3d.statistics(cc, no_slice_conversion=True)\n\n        if node_type == \"centroid\":\n            ys_t = cc_stats[\"centroids\"][1:, 0]\n        else:\n            cc_range = cc_stats[\"bounding_boxes\"][1:, :2].astype(float)\n            ys_t = cc_range[\n                np.arange(cc_range.shape[0]), \n                np.argmax(np.abs(cc_range - y_center), axis=1)\n            ]\n\n        ys_t = ys_t[np.abs(ys_t - y_center) &lt;= threshold]\n\n        if ys_t.shape[0] == 0:\n            all_curr_states.append(np.asarray([]))\n            all_prev_indexes.append(np.asarray([]))\n            all_costs.append(np.asarray([]))\n            continue\n\n        all_curr_states.append(ys_t)\n        if x == 0:\n            all_prev_indexes.append(np.zeros_like(ys_t, dtype=np.int64))\n            all_costs.append(np.zeros_like(ys_t))\n            continue\n\n        a = 1\n        while True:\n            x_tm1 = x - a\n            if len(all_curr_states[x_tm1]) &gt; 0: break\n            a += 1\n        all_ys_tm1 = all_curr_states[x_tm1]\n        all_cost_tm1 = all_costs[x_tm1]\n        cost = alpha * dist_cost(ys_t[:, None], all_ys_tm1[None, :])\n\n        x_tm2 = x_tm1 - 1\n        all_ys_tm2 = all_curr_states[x_tm2]\n        if x &gt; 1 and alpha &lt; 1 and all_ys_tm2.shape[0]:\n            smooth_cost = smoothness_cost(\n                ys_t[:, None, None], \n                all_ys_tm1[None, :, None], \n                all_ys_tm2[None, None, :]\n            )\n            cost = cost[:, :, None] + smooth_cost\n            cost = np.min(cost, axis=2)\n\n        best_indexes = np.argmin(cost, axis=1)\n        all_prev_indexes.append(best_indexes)\n        all_costs.append(all_cost_tm1[best_indexes] + np.min(cost, axis=1))\n\n    path = np.zeros((W, ))\n    j = None\n    for x in range(W, -1, -1):\n        if x == W  or j is None:\n            costs = all_costs[x-1]\n            if costs.shape[0] == 0:\n                continue\n            j = np.argmin(costs)\n            continue\n\n        s = all_curr_states[x]\n        pidx = all_prev_indexes[x]\n        if len(s) == 0:\n            path[x] = np.nan\n            continue\n        path[x] = s[j].item()\n        j = pidx[j]\n\n    path = interpolate_nan(path)\n    y_mm_traces = px2mm_along_y(y_center - path)\n    y_mv_traces = mm2mv(y_mm_traces)\n    return y_mv_traces\n</code></pre>\n<p>connected components analysis is done with the cc3d library</p>\n<p>You will notice here that there's the option to use centroids and the option to use farthest, centroids = center of connected 1s, farthest = the point in the connected ones farthest from the baseline of that row. Based on experience, it seems \"farthest\" works best because it is more capable of detecting peaks than \"centroid\" option.</p>\n<p>How to use:</p>\n<pre><code>r_trace2 = viterbi(\n        signals, \n        row_mm_center=89+(36 * row_idx), \n        row_mm_spacing=36,\n        vpx_spacing=vpx_spacing, \n        hpx_spacing=hpx_spacing,\n        alpha=0.5,\n        mm_tol=0,\n        node_type=\"farthest\",\n    )\n</code></pre>\n<p>signals is your binary mask of size (H, W) without cropping, containing the full signal segments in the 4 rows, like this:</p>\n<p><code>row_mm_center</code> is the baseline of a given row in mm. the first row seems to have a baseline of 89mm, and each baseline consecutive is 36mm apart, so you get the gist of what <code>row_idx</code> is. <code>vpx_spacing</code> and <code>hpx_spacing</code> are just pixel spacings of the major gridlines, its something around 39 to 40 pixels, representing 5mm (they are not hardcoded but calculated automatically, but that's the range) for an image resolution of 1720 x 2200. Alpha is the weight value for the cost function, mm_tol is a tolerance level of sorts, it kind of permits the algorithm to look into centroids or farthest points of contiguous 1s that are farther than half of 36mm from the baseline for the given row.</p>\n<p>So, with an input image of this:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5338718%2F5199eef20f81efd82bf0c6d4f85dcbbe%2Fsig.jpg?generation=1766010536134554&amp;alt=media\" alt=\"\"></p>\n<p>and a <code>row_idx</code> of 0</p>\n<p>after performing the necessary upsampling to match sample frequency and adjusting for DC (baseline) correction, you will get exactly this:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5338718%2Fd0d9859f4b9fe894c53334e7aa3284a4%2FScreenshot%202025-12-17%20232754.png?generation=1766010591762652&amp;alt=media\" alt=\"\"></p>\n<p>And <code>row_idx = 1</code> will give you the trace of the second row, and <code>row_idx = 2</code> the trace of the third row and so on…</p>\n<p>As for the cost function I talked about earlier, it is a combination of absolute distance and smoothness as seen in the toy program. </p>\n<p>If you are wondering why smoothness is computed that way instead of using the values from t-2 that best optimize the path leading to the centroids at t-1 instead, this is why (Took me a while to wrap my head around):</p>\n<p>Smoothness is not a first order cost, since it depends on the last 2 timesteps and current timesteps, so the formula |(y_t - y_{t-1}) - (y_{t-1}- y_{t-2})|</p>\n<p>y_t is a tensor of size n and ytm1 is a tensor of size m, now instead of y_{t-2} to be the previous best centroids that leads to each centroid in y_{t-1} (which would make it of size m), we should just get all the centroids at t-2 instead (all_curr_states[x-2]) which will be of size k, and m =/= n =/= k (at least we can assume so because it is not necessarily meant to be)</p>\n<p>now these shapes are obviously incompatible (n, m, k), but the idea is to compute the smoothness cost, taken into account each triple combinations in the size m, n and k centroids arrays .</p>\n<p>So y_t can be expanded to shape (n, 1, 1), then y_{t-1} can be expanded to shape (1, m, 1) and y_{t-2}will be (1, 1, k), this way, every shape the arrays broadcast to one another to give us a cost array of size (n, m, k), then add this to the distance cost of shape (n, m, 1) and compute the minimum along the last 2 axes. Hope that explains it.</p>\n<p><strong>PS: This function does need some improvements and is not meant to be used blindly as is, its just for demonstration. If you or anyone else have suggestions, do share them, I too am looking for an even better version or alternative.</strong></p>",
      "rawMarkdown": "I wouldn't say this is best (or even good) practice or a standard pipeline, however, something that can work is the Viterbi's algorithm, although this is not an algorithm you'd normally use here, it more likely has applications elsewhere (NLP and error correction in telecom signals)\n\nThe idea is this:\n\nyou have a binary mask of size (H x W), with 0s corresponding to background and 1s corresponding to foreground (your signal segment). For each column along the second axis (axis=1), find the centroid (or farthest point from baseline) of contiguous 1s, you will use these centroids to construct a DAG (Directed Acyclic Graph), with the centroids being the nodes and connection between centroids of consecutive steps being the edges, but you will not store all of the graph nodes and edges, rather you do this:\n\n1: For each centroid at position t, associate it with all centroids at t-1 and compute a cost function (I will explain a 'good enough' cost function later). EG: suppose position t has 4 contiguous 1s, that is 4 centroids, if position t-1 had 3 centroids, now you will match each centroid at t to all centroids at t-1, forming 12 edges or paths.\n\n2: For each of the centroids in t, get the centroid from t-1 that best optimizes the cost function, in other words, get the edge with the least cost for a given centroid. That will be the centroid from the previous timestep that leads to the current timestep, you store the centroids at timestep t, as well as their corresponding best previous centroids from each timestep and the corresponding cost values. Do note that the cost values from the previous timesteps are propagated into the future by adding them to the currently computed cost, EG: cost_t = cost_{t-1} + D(y_t, y_{t-1}), where D is our base cost.\n\n3: continue this process until you get to the very last timestep t=T-1\n\n4: At t=T-1, you backtrack. You do this by getting the centroid with the least cost value at t=T-1, let's call that y_{T-1}, then you retrieve the best centroid from T-2 that lead to y_{T-1}, which will be y_{T-2}, then you retrieve the best centroid from T-3 that lead to y_{T-2}, which is y_{T-3}, you do this until you get to t=0. And that is the Viterbi's solution to this problem. It selects the path of least resistance (or least cost if you may), This is somewhat equivalent to constructing the full graph and running an A-star algorithm to find the best path, but it is far more efficient.\n\nThis is an example code that still needs a lot of improvement:\n\n```\ndef interpolate_nan(array: np.ndarray, inplace: bool=False) -> np.ndarray:\n    if not inplace:\n        array = array.copy()\n    nans = np.isnan(array)\n    indices = np.arange(len(array))\n    xp = indices[not_nans := ~nans] \n    fp = array[not_nans]\n    array[nans] = np.interp(indices[nans], xp, fp)\n    return array\n\ndef viterbi(\n        mask: np.ndarray, \n        row_mm_center: float,\n        row_mm_spacing: float,\n        hpx_spacing: float,\n        vpx_spacing: float,\n        vmm_spacing: float=5,\n        hcrop_offset: int=8,\n        mm_per_mv: float=10,\n        alpha: float=0.5,\n        mm_tol: float=0.0,\n        node_type: Literal[\"centroid\", \"farthest\"]=\"centroid\"\n    ):\n    assert node_type in [\"centroid\", \"farthest\"]\n    assert alpha >= 0 and alpha <= 1 \n\n    mm2px_along_y = lambda mm : mm * (vpx_spacing / vmm_spacing)\n    px2mm_along_y = lambda px : px * (vmm_spacing / vpx_spacing)\n    mm2mv = lambda mm : mm / mm_per_mv\n    dist_cost = lambda x1, x2: np.abs(x1 - x2)\n    smoothness_cost = lambda xt, xtm1, xtm2: np.abs((xt - xtm1) - (xtm1 - xtm2))\n    \n    y_center = mm2px_along_y(row_mm_center)\n    row_spacing = mm2px_along_y(row_mm_spacing)\n    tol = mm2px_along_y(mm_tol)\n    threshold = row_spacing / 2 + tol\n\n    _, xs = np.where(mask)\n    mask = mask[:, xs.min()+int(hpx_spacing)+hcrop_offset : xs.max()]\n    W = mask.shape[1]\n\n    all_curr_states = []\n    all_prev_indexes = []\n    all_costs = []\n\n    for x in range(0, W):\n        m = mask[:, x]\n        if np.where(m)[0].shape[0] == 0:\n            all_curr_states.append(np.asarray([]))\n            all_prev_indexes.append(np.asarray([]))\n            all_costs.append(np.asarray([]))\n            continue\n\n        cc = cc3d.connected_components(m[:, None])\n        cc_stats = cc3d.statistics(cc, no_slice_conversion=True)\n\n        if node_type == \"centroid\":\n            ys_t = cc_stats[\"centroids\"][1:, 0]\n        else:\n            cc_range = cc_stats[\"bounding_boxes\"][1:, :2].astype(float)\n            ys_t = cc_range[\n                np.arange(cc_range.shape[0]), \n                np.argmax(np.abs(cc_range - y_center), axis=1)\n            ]\n        \n        ys_t = ys_t[np.abs(ys_t - y_center) <= threshold]\n\n        if ys_t.shape[0] == 0:\n            all_curr_states.append(np.asarray([]))\n            all_prev_indexes.append(np.asarray([]))\n            all_costs.append(np.asarray([]))\n            continue\n        \n        all_curr_states.append(ys_t)\n        if x == 0:\n            all_prev_indexes.append(np.zeros_like(ys_t, dtype=np.int64))\n            all_costs.append(np.zeros_like(ys_t))\n            continue\n\n        a = 1\n        while True:\n            x_tm1 = x - a\n            if len(all_curr_states[x_tm1]) > 0: break\n            a += 1\n        all_ys_tm1 = all_curr_states[x_tm1]\n        all_cost_tm1 = all_costs[x_tm1]\n        cost = alpha * dist_cost(ys_t[:, None], all_ys_tm1[None, :])\n\n        x_tm2 = x_tm1 - 1\n        all_ys_tm2 = all_curr_states[x_tm2]\n        if x > 1 and alpha < 1 and all_ys_tm2.shape[0]:\n            smooth_cost = smoothness_cost(\n                ys_t[:, None, None], \n                all_ys_tm1[None, :, None], \n                all_ys_tm2[None, None, :]\n            )\n            cost = cost[:, :, None] + smooth_cost\n            cost = np.min(cost, axis=2)\n\n        best_indexes = np.argmin(cost, axis=1)\n        all_prev_indexes.append(best_indexes)\n        all_costs.append(all_cost_tm1[best_indexes] + np.min(cost, axis=1))\n\n    path = np.zeros((W, ))\n    j = None\n    for x in range(W, -1, -1):\n        if x == W  or j is None:\n            costs = all_costs[x-1]\n            if costs.shape[0] == 0:\n                continue\n            j = np.argmin(costs)\n            continue\n            \n        s = all_curr_states[x]\n        pidx = all_prev_indexes[x]\n        if len(s) == 0:\n            path[x] = np.nan\n            continue\n        path[x] = s[j].item()\n        j = pidx[j]\n\n    path = interpolate_nan(path)\n    y_mm_traces = px2mm_along_y(y_center - path)\n    y_mv_traces = mm2mv(y_mm_traces)\n    return y_mv_traces\n```\n\nconnected components analysis is done with the cc3d library\n\nYou will notice here that there's the option to use centroids and the option to use farthest, centroids = center of connected 1s, farthest = the point in the connected ones farthest from the baseline of that row. Based on experience, it seems \"farthest\" works best because it is more capable of detecting peaks than \"centroid\" option.\n\nHow to use:\n\n```\nr_trace2 = viterbi(\n        signals, \n        row_mm_center=89+(36 * row_idx), \n        row_mm_spacing=36,\n        vpx_spacing=vpx_spacing, \n        hpx_spacing=hpx_spacing,\n        alpha=0.5,\n        mm_tol=0,\n        node_type=\"farthest\",\n    )\n```\nsignals is your binary mask of size (H, W) without cropping, containing the full signal segments in the 4 rows, like this:\n\n`row_mm_center` is the baseline of a given row in mm. the first row seems to have a baseline of 89mm, and each baseline consecutive is 36mm apart, so you get the gist of what `row_idx` is. `vpx_spacing` and `hpx_spacing` are just pixel spacings of the major gridlines, its something around 39 to 40 pixels, representing 5mm (they are not hardcoded but calculated automatically, but that's the range) for an image resolution of 1720 x 2200. Alpha is the weight value for the cost function, mm_tol is a tolerance level of sorts, it kind of permits the algorithm to look into centroids or farthest points of contiguous 1s that are farther than half of 36mm from the baseline for the given row.\n\n\nSo, with an input image of this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5338718%2F5199eef20f81efd82bf0c6d4f85dcbbe%2Fsig.jpg?generation=1766010536134554&alt=media)\n\nand a `row_idx` of 0\n\nafter performing the necessary upsampling to match sample frequency and adjusting for DC (baseline) correction, you will get exactly this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5338718%2Fd0d9859f4b9fe894c53334e7aa3284a4%2FScreenshot%202025-12-17%20232754.png?generation=1766010591762652&alt=media)\n\nAnd `row_idx = 1` will give you the trace of the second row, and `row_idx = 2` the trace of the third row and so on...\n\nAs for the cost function I talked about earlier, it is a combination of absolute distance and smoothness as seen in the toy program. \n\nIf you are wondering why smoothness is computed that way instead of using the values from t-2 that best optimize the path leading to the centroids at t-1 instead, this is why (Took me a while to wrap my head around):\n\nSmoothness is not a first order cost, since it depends on the last 2 timesteps and current timesteps, so the formula |(y_t - y_{t-1}) - (y_{t-1}- y_{t-2})|\n\ny_t is a tensor of size n and ytm1 is a tensor of size m, now instead of y_{t-2} to be the previous best centroids that leads to each centroid in y_{t-1} (which would make it of size m), we should just get all the centroids at t-2 instead (all_curr_states[x-2]) which will be of size k, and m =/= n =/= k (at least we can assume so because it is not necessarily meant to be)\n\nnow these shapes are obviously incompatible (n, m, k), but the idea is to compute the smoothness cost, taken into account each triple combinations in the size m, n and k centroids arrays .\n\nSo y_t can be expanded to shape (n, 1, 1), then y_{t-1} can be expanded to shape (1, m, 1) and y_{t-2}will be (1, 1, k), this way, every shape the arrays broadcast to one another to give us a cost array of size (n, m, k), then add this to the distance cost of shape (n, m, 1) and compute the minimum along the last 2 axes. Hope that explains it.\n\n**PS: This function does need some improvements and is not meant to be used blindly as is, its just for demonstration. If you or anyone else have suggestions, do share them, I too am looking for an even better version or alternative.**",
      "votes": null
    },
    {
      "id": "3378368",
      "postDate": "12/18/2025 00:44:30",
      "content": "<p>This is a genuinely insightful contribution thank you for taking the time to explain both the idea and the reasoning behind it so clearly.</p>",
      "rawMarkdown": "This is a genuinely insightful contribution thank you for taking the time to explain both the idea and the reasoning behind it so clearly.",
      "votes": null
    },
    {
      "id": "3378376",
      "postDate": "12/18/2025 01:33:25",
      "content": "<p>You're welcome.</p>",
      "rawMarkdown": "You're welcome.",
      "votes": null
    },
    {
      "id": "3378480",
      "postDate": "12/18/2025 10:00:31",
      "content": "<p>Thank you for the explanation!!!\n I will use it as a reference!</p>",
      "rawMarkdown": "Thank you for the explanation!!!\n I will use it as a reference!",
      "votes": null
    },
    {
      "id": "3379734",
      "postDate": "12/20/2025 15:19:21",
      "content": "<p>Thanks bro this will help me a lot</p>",
      "rawMarkdown": "Thanks bro this will help me a lot",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3378328,
      "author_name": "henrychibueze",
      "author_url": "",
      "post_date": "12/17/2025 22:33:17",
      "content": "<p>I wouldn't say this is best (or even good) practice or a standard pipeline, however, something that can work is the Viterbi's algorithm, although this is not an algorithm you'd normally use here, it more likely has applications elsewhere (NLP and error correction in telecom signals)</p>\n<p>The idea is this:</p>\n<p>you have a binary mask of size (H x W), with 0s corresponding to background and 1s corresponding to foreground (your signal segment). For each column along the second axis (axis=1), find the centroid (or farthest point from baseline) of contiguous 1s, you will use these centroids to construct a DAG (Directed Acyclic Graph), with the centroids being the nodes and connection between centroids of consecutive steps being the edges, but you will not store all of the graph nodes and edges, rather you do this:</p>\n<p>1: For each centroid at position t, associate it with all centroids at t-1 and compute a cost function (I will explain a 'good enough' cost function later). EG: suppose position t has 4 contiguous 1s, that is 4 centroids, if position t-1 had 3 centroids, now you will match each centroid at t to all centroids at t-1, forming 12 edges or paths.</p>\n<p>2: For each of the centroids in t, get the centroid from t-1 that best optimizes the cost function, in other words, get the edge with the least cost for a given centroid. That will be the centroid from the previous timestep that leads to the current timestep, you store the centroids at timestep t, as well as their corresponding best previous centroids from each timestep and the corresponding cost values. Do note that the cost values from the previous timesteps are propagated into the future by adding them to the currently computed cost, EG: cost_t = cost_{t-1} + D(y_t, y_{t-1}), where D is our base cost.</p>\n<p>3: continue this process until you get to the very last timestep t=T-1</p>\n<p>4: At t=T-1, you backtrack. You do this by getting the centroid with the least cost value at t=T-1, let's call that y_{T-1}, then you retrieve the best centroid from T-2 that lead to y_{T-1}, which will be y_{T-2}, then you retrieve the best centroid from T-3 that lead to y_{T-2}, which is y_{T-3}, you do this until you get to t=0. And that is the Viterbi's solution to this problem. It selects the path of least resistance (or least cost if you may), This is somewhat equivalent to constructing the full graph and running an A-star algorithm to find the best path, but it is far more efficient.</p>\n<p>This is an example code that still needs a lot of improvement:</p>\n<pre><code>def interpolate_nan(array: np.ndarray, inplace: bool=False) -&gt; np.ndarray:\n    if not inplace:\n        array = array.copy()\n    nans = np.isnan(array)\n    indices = np.arange(len(array))\n    xp = indices[not_nans := ~nans] \n    fp = array[not_nans]\n    array[nans] = np.interp(indices[nans], xp, fp)\n    return array\n\ndef viterbi(\n        mask: np.ndarray, \n        row_mm_center: float,\n        row_mm_spacing: float,\n        hpx_spacing: float,\n        vpx_spacing: float,\n        vmm_spacing: float=5,\n        hcrop_offset: int=8,\n        mm_per_mv: float=10,\n        alpha: float=0.5,\n        mm_tol: float=0.0,\n        node_type: Literal[\"centroid\", \"farthest\"]=\"centroid\"\n    ):\n    assert node_type in [\"centroid\", \"farthest\"]\n    assert alpha &gt;= 0 and alpha &lt;= 1 \n\n    mm2px_along_y = lambda mm : mm * (vpx_spacing / vmm_spacing)\n    px2mm_along_y = lambda px : px * (vmm_spacing / vpx_spacing)\n    mm2mv = lambda mm : mm / mm_per_mv\n    dist_cost = lambda x1, x2: np.abs(x1 - x2)\n    smoothness_cost = lambda xt, xtm1, xtm2: np.abs((xt - xtm1) - (xtm1 - xtm2))\n\n    y_center = mm2px_along_y(row_mm_center)\n    row_spacing = mm2px_along_y(row_mm_spacing)\n    tol = mm2px_along_y(mm_tol)\n    threshold = row_spacing / 2 + tol\n\n    _, xs = np.where(mask)\n    mask = mask[:, xs.min()+int(hpx_spacing)+hcrop_offset : xs.max()]\n    W = mask.shape[1]\n\n    all_curr_states = []\n    all_prev_indexes = []\n    all_costs = []\n\n    for x in range(0, W):\n        m = mask[:, x]\n        if np.where(m)[0].shape[0] == 0:\n            all_curr_states.append(np.asarray([]))\n            all_prev_indexes.append(np.asarray([]))\n            all_costs.append(np.asarray([]))\n            continue\n\n        cc = cc3d.connected_components(m[:, None])\n        cc_stats = cc3d.statistics(cc, no_slice_conversion=True)\n\n        if node_type == \"centroid\":\n            ys_t = cc_stats[\"centroids\"][1:, 0]\n        else:\n            cc_range = cc_stats[\"bounding_boxes\"][1:, :2].astype(float)\n            ys_t = cc_range[\n                np.arange(cc_range.shape[0]), \n                np.argmax(np.abs(cc_range - y_center), axis=1)\n            ]\n\n        ys_t = ys_t[np.abs(ys_t - y_center) &lt;= threshold]\n\n        if ys_t.shape[0] == 0:\n            all_curr_states.append(np.asarray([]))\n            all_prev_indexes.append(np.asarray([]))\n            all_costs.append(np.asarray([]))\n            continue\n\n        all_curr_states.append(ys_t)\n        if x == 0:\n            all_prev_indexes.append(np.zeros_like(ys_t, dtype=np.int64))\n            all_costs.append(np.zeros_like(ys_t))\n            continue\n\n        a = 1\n        while True:\n            x_tm1 = x - a\n            if len(all_curr_states[x_tm1]) &gt; 0: break\n            a += 1\n        all_ys_tm1 = all_curr_states[x_tm1]\n        all_cost_tm1 = all_costs[x_tm1]\n        cost = alpha * dist_cost(ys_t[:, None], all_ys_tm1[None, :])\n\n        x_tm2 = x_tm1 - 1\n        all_ys_tm2 = all_curr_states[x_tm2]\n        if x &gt; 1 and alpha &lt; 1 and all_ys_tm2.shape[0]:\n            smooth_cost = smoothness_cost(\n                ys_t[:, None, None], \n                all_ys_tm1[None, :, None], \n                all_ys_tm2[None, None, :]\n            )\n            cost = cost[:, :, None] + smooth_cost\n            cost = np.min(cost, axis=2)\n\n        best_indexes = np.argmin(cost, axis=1)\n        all_prev_indexes.append(best_indexes)\n        all_costs.append(all_cost_tm1[best_indexes] + np.min(cost, axis=1))\n\n    path = np.zeros((W, ))\n    j = None\n    for x in range(W, -1, -1):\n        if x == W  or j is None:\n            costs = all_costs[x-1]\n            if costs.shape[0] == 0:\n                continue\n            j = np.argmin(costs)\n            continue\n\n        s = all_curr_states[x]\n        pidx = all_prev_indexes[x]\n        if len(s) == 0:\n            path[x] = np.nan\n            continue\n        path[x] = s[j].item()\n        j = pidx[j]\n\n    path = interpolate_nan(path)\n    y_mm_traces = px2mm_along_y(y_center - path)\n    y_mv_traces = mm2mv(y_mm_traces)\n    return y_mv_traces\n</code></pre>\n<p>connected components analysis is done with the cc3d library</p>\n<p>You will notice here that there's the option to use centroids and the option to use farthest, centroids = center of connected 1s, farthest = the point in the connected ones farthest from the baseline of that row. Based on experience, it seems \"farthest\" works best because it is more capable of detecting peaks than \"centroid\" option.</p>\n<p>How to use:</p>\n<pre><code>r_trace2 = viterbi(\n        signals, \n        row_mm_center=89+(36 * row_idx), \n        row_mm_spacing=36,\n        vpx_spacing=vpx_spacing, \n        hpx_spacing=hpx_spacing,\n        alpha=0.5,\n        mm_tol=0,\n        node_type=\"farthest\",\n    )\n</code></pre>\n<p>signals is your binary mask of size (H, W) without cropping, containing the full signal segments in the 4 rows, like this:</p>\n<p><code>row_mm_center</code> is the baseline of a given row in mm. the first row seems to have a baseline of 89mm, and each baseline consecutive is 36mm apart, so you get the gist of what <code>row_idx</code> is. <code>vpx_spacing</code> and <code>hpx_spacing</code> are just pixel spacings of the major gridlines, its something around 39 to 40 pixels, representing 5mm (they are not hardcoded but calculated automatically, but that's the range) for an image resolution of 1720 x 2200. Alpha is the weight value for the cost function, mm_tol is a tolerance level of sorts, it kind of permits the algorithm to look into centroids or farthest points of contiguous 1s that are farther than half of 36mm from the baseline for the given row.</p>\n<p>So, with an input image of this:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5338718%2F5199eef20f81efd82bf0c6d4f85dcbbe%2Fsig.jpg?generation=1766010536134554&amp;alt=media\" alt=\"\"></p>\n<p>and a <code>row_idx</code> of 0</p>\n<p>after performing the necessary upsampling to match sample frequency and adjusting for DC (baseline) correction, you will get exactly this:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5338718%2Fd0d9859f4b9fe894c53334e7aa3284a4%2FScreenshot%202025-12-17%20232754.png?generation=1766010591762652&amp;alt=media\" alt=\"\"></p>\n<p>And <code>row_idx = 1</code> will give you the trace of the second row, and <code>row_idx = 2</code> the trace of the third row and so on…</p>\n<p>As for the cost function I talked about earlier, it is a combination of absolute distance and smoothness as seen in the toy program. </p>\n<p>If you are wondering why smoothness is computed that way instead of using the values from t-2 that best optimize the path leading to the centroids at t-1 instead, this is why (Took me a while to wrap my head around):</p>\n<p>Smoothness is not a first order cost, since it depends on the last 2 timesteps and current timesteps, so the formula |(y_t - y_{t-1}) - (y_{t-1}- y_{t-2})|</p>\n<p>y_t is a tensor of size n and ytm1 is a tensor of size m, now instead of y_{t-2} to be the previous best centroids that leads to each centroid in y_{t-1} (which would make it of size m), we should just get all the centroids at t-2 instead (all_curr_states[x-2]) which will be of size k, and m =/= n =/= k (at least we can assume so because it is not necessarily meant to be)</p>\n<p>now these shapes are obviously incompatible (n, m, k), but the idea is to compute the smoothness cost, taken into account each triple combinations in the size m, n and k centroids arrays .</p>\n<p>So y_t can be expanded to shape (n, 1, 1), then y_{t-1} can be expanded to shape (1, m, 1) and y_{t-2}will be (1, 1, k), this way, every shape the arrays broadcast to one another to give us a cost array of size (n, m, k), then add this to the distance cost of shape (n, m, 1) and compute the minimum along the last 2 axes. Hope that explains it.</p>\n<p><strong>PS: This function does need some improvements and is not meant to be used blindly as is, its just for demonstration. If you or anyone else have suggestions, do share them, I too am looking for an even better version or alternative.</strong></p>",
      "votes": null,
      "replies": [
        {
          "id": 3378480,
          "author_name": "nakaosyoya",
          "author_url": "",
          "post_date": "12/18/2025 10:00:31",
          "content": "<p>Thank you for the explanation!!!\n I will use it as a reference!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3378368,
      "author_name": "ladiposamson",
      "author_url": "",
      "post_date": "12/18/2025 00:44:30",
      "content": "<p>This is a genuinely insightful contribution thank you for taking the time to explain both the idea and the reasoning behind it so clearly.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3378376,
          "author_name": "henrychibueze",
          "author_url": "",
          "post_date": "12/18/2025 01:33:25",
          "content": "<p>You're welcome.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3379734,
      "author_name": "craftycode",
      "author_url": "",
      "post_date": "12/20/2025 15:19:21",
      "content": "<p>Thanks bro this will help me a lot</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3376801": "I am trying to digitize ECG waveforms from images after extracting the waveform region using a neural network (segmentation model).\n\nWhile I can obtain a binary mask or a cleaned waveform image from the model, I am not sure about the correct or standard way to convert this image-based waveform into a 1D digital signal (time series).\n\nI would really appreciate it if anyone could share a recommended pipeline or best practices for digitizing waveforms from ECG images after NN-based segmentation.\n\nThank you very much!",
    "3378328": "I wouldn't say this is best (or even good) practice or a standard pipeline, however, something that can work is the Viterbi's algorithm, although this is not an algorithm you'd normally use here, it more likely has applications elsewhere (NLP and error correction in telecom signals)\n\nThe idea is this:\n\nyou have a binary mask of size (H x W), with 0s corresponding to background and 1s corresponding to foreground (your signal segment). For each column along the second axis (axis=1), find the centroid (or farthest point from baseline) of contiguous 1s, you will use these centroids to construct a DAG (Directed Acyclic Graph), with the centroids being the nodes and connection between centroids of consecutive steps being the edges, but you will not store all of the graph nodes and edges, rather you do this:\n\n1: For each centroid at position t, associate it with all centroids at t-1 and compute a cost function (I will explain a 'good enough' cost function later). EG: suppose position t has 4 contiguous 1s, that is 4 centroids, if position t-1 had 3 centroids, now you will match each centroid at t to all centroids at t-1, forming 12 edges or paths.\n\n2: For each of the centroids in t, get the centroid from t-1 that best optimizes the cost function, in other words, get the edge with the least cost for a given centroid. That will be the centroid from the previous timestep that leads to the current timestep, you store the centroids at timestep t, as well as their corresponding best previous centroids from each timestep and the corresponding cost values. Do note that the cost values from the previous timesteps are propagated into the future by adding them to the currently computed cost, EG: cost_t = cost_{t-1} + D(y_t, y_{t-1}), where D is our base cost.\n\n3: continue this process until you get to the very last timestep t=T-1\n\n4: At t=T-1, you backtrack. You do this by getting the centroid with the least cost value at t=T-1, let's call that y_{T-1}, then you retrieve the best centroid from T-2 that lead to y_{T-1}, which will be y_{T-2}, then you retrieve the best centroid from T-3 that lead to y_{T-2}, which is y_{T-3}, you do this until you get to t=0. And that is the Viterbi's solution to this problem. It selects the path of least resistance (or least cost if you may), This is somewhat equivalent to constructing the full graph and running an A-star algorithm to find the best path, but it is far more efficient.\n\nThis is an example code that still needs a lot of improvement:\n\n```\ndef interpolate_nan(array: np.ndarray, inplace: bool=False) -> np.ndarray:\n    if not inplace:\n        array = array.copy()\n    nans = np.isnan(array)\n    indices = np.arange(len(array))\n    xp = indices[not_nans := ~nans] \n    fp = array[not_nans]\n    array[nans] = np.interp(indices[nans], xp, fp)\n    return array\n\ndef viterbi(\n        mask: np.ndarray, \n        row_mm_center: float,\n        row_mm_spacing: float,\n        hpx_spacing: float,\n        vpx_spacing: float,\n        vmm_spacing: float=5,\n        hcrop_offset: int=8,\n        mm_per_mv: float=10,\n        alpha: float=0.5,\n        mm_tol: float=0.0,\n        node_type: Literal[\"centroid\", \"farthest\"]=\"centroid\"\n    ):\n    assert node_type in [\"centroid\", \"farthest\"]\n    assert alpha >= 0 and alpha <= 1 \n\n    mm2px_along_y = lambda mm : mm * (vpx_spacing / vmm_spacing)\n    px2mm_along_y = lambda px : px * (vmm_spacing / vpx_spacing)\n    mm2mv = lambda mm : mm / mm_per_mv\n    dist_cost = lambda x1, x2: np.abs(x1 - x2)\n    smoothness_cost = lambda xt, xtm1, xtm2: np.abs((xt - xtm1) - (xtm1 - xtm2))\n    \n    y_center = mm2px_along_y(row_mm_center)\n    row_spacing = mm2px_along_y(row_mm_spacing)\n    tol = mm2px_along_y(mm_tol)\n    threshold = row_spacing / 2 + tol\n\n    _, xs = np.where(mask)\n    mask = mask[:, xs.min()+int(hpx_spacing)+hcrop_offset : xs.max()]\n    W = mask.shape[1]\n\n    all_curr_states = []\n    all_prev_indexes = []\n    all_costs = []\n\n    for x in range(0, W):\n        m = mask[:, x]\n        if np.where(m)[0].shape[0] == 0:\n            all_curr_states.append(np.asarray([]))\n            all_prev_indexes.append(np.asarray([]))\n            all_costs.append(np.asarray([]))\n            continue\n\n        cc = cc3d.connected_components(m[:, None])\n        cc_stats = cc3d.statistics(cc, no_slice_conversion=True)\n\n        if node_type == \"centroid\":\n            ys_t = cc_stats[\"centroids\"][1:, 0]\n        else:\n            cc_range = cc_stats[\"bounding_boxes\"][1:, :2].astype(float)\n            ys_t = cc_range[\n                np.arange(cc_range.shape[0]), \n                np.argmax(np.abs(cc_range - y_center), axis=1)\n            ]\n        \n        ys_t = ys_t[np.abs(ys_t - y_center) <= threshold]\n\n        if ys_t.shape[0] == 0:\n            all_curr_states.append(np.asarray([]))\n            all_prev_indexes.append(np.asarray([]))\n            all_costs.append(np.asarray([]))\n            continue\n        \n        all_curr_states.append(ys_t)\n        if x == 0:\n            all_prev_indexes.append(np.zeros_like(ys_t, dtype=np.int64))\n            all_costs.append(np.zeros_like(ys_t))\n            continue\n\n        a = 1\n        while True:\n            x_tm1 = x - a\n            if len(all_curr_states[x_tm1]) > 0: break\n            a += 1\n        all_ys_tm1 = all_curr_states[x_tm1]\n        all_cost_tm1 = all_costs[x_tm1]\n        cost = alpha * dist_cost(ys_t[:, None], all_ys_tm1[None, :])\n\n        x_tm2 = x_tm1 - 1\n        all_ys_tm2 = all_curr_states[x_tm2]\n        if x > 1 and alpha < 1 and all_ys_tm2.shape[0]:\n            smooth_cost = smoothness_cost(\n                ys_t[:, None, None], \n                all_ys_tm1[None, :, None], \n                all_ys_tm2[None, None, :]\n            )\n            cost = cost[:, :, None] + smooth_cost\n            cost = np.min(cost, axis=2)\n\n        best_indexes = np.argmin(cost, axis=1)\n        all_prev_indexes.append(best_indexes)\n        all_costs.append(all_cost_tm1[best_indexes] + np.min(cost, axis=1))\n\n    path = np.zeros((W, ))\n    j = None\n    for x in range(W, -1, -1):\n        if x == W  or j is None:\n            costs = all_costs[x-1]\n            if costs.shape[0] == 0:\n                continue\n            j = np.argmin(costs)\n            continue\n            \n        s = all_curr_states[x]\n        pidx = all_prev_indexes[x]\n        if len(s) == 0:\n            path[x] = np.nan\n            continue\n        path[x] = s[j].item()\n        j = pidx[j]\n\n    path = interpolate_nan(path)\n    y_mm_traces = px2mm_along_y(y_center - path)\n    y_mv_traces = mm2mv(y_mm_traces)\n    return y_mv_traces\n```\n\nconnected components analysis is done with the cc3d library\n\nYou will notice here that there's the option to use centroids and the option to use farthest, centroids = center of connected 1s, farthest = the point in the connected ones farthest from the baseline of that row. Based on experience, it seems \"farthest\" works best because it is more capable of detecting peaks than \"centroid\" option.\n\nHow to use:\n\n```\nr_trace2 = viterbi(\n        signals, \n        row_mm_center=89+(36 * row_idx), \n        row_mm_spacing=36,\n        vpx_spacing=vpx_spacing, \n        hpx_spacing=hpx_spacing,\n        alpha=0.5,\n        mm_tol=0,\n        node_type=\"farthest\",\n    )\n```\nsignals is your binary mask of size (H, W) without cropping, containing the full signal segments in the 4 rows, like this:\n\n`row_mm_center` is the baseline of a given row in mm. the first row seems to have a baseline of 89mm, and each baseline consecutive is 36mm apart, so you get the gist of what `row_idx` is. `vpx_spacing` and `hpx_spacing` are just pixel spacings of the major gridlines, its something around 39 to 40 pixels, representing 5mm (they are not hardcoded but calculated automatically, but that's the range) for an image resolution of 1720 x 2200. Alpha is the weight value for the cost function, mm_tol is a tolerance level of sorts, it kind of permits the algorithm to look into centroids or farthest points of contiguous 1s that are farther than half of 36mm from the baseline for the given row.\n\n\nSo, with an input image of this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5338718%2F5199eef20f81efd82bf0c6d4f85dcbbe%2Fsig.jpg?generation=1766010536134554&alt=media)\n\nand a `row_idx` of 0\n\nafter performing the necessary upsampling to match sample frequency and adjusting for DC (baseline) correction, you will get exactly this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5338718%2Fd0d9859f4b9fe894c53334e7aa3284a4%2FScreenshot%202025-12-17%20232754.png?generation=1766010591762652&alt=media)\n\nAnd `row_idx = 1` will give you the trace of the second row, and `row_idx = 2` the trace of the third row and so on...\n\nAs for the cost function I talked about earlier, it is a combination of absolute distance and smoothness as seen in the toy program. \n\nIf you are wondering why smoothness is computed that way instead of using the values from t-2 that best optimize the path leading to the centroids at t-1 instead, this is why (Took me a while to wrap my head around):\n\nSmoothness is not a first order cost, since it depends on the last 2 timesteps and current timesteps, so the formula |(y_t - y_{t-1}) - (y_{t-1}- y_{t-2})|\n\ny_t is a tensor of size n and ytm1 is a tensor of size m, now instead of y_{t-2} to be the previous best centroids that leads to each centroid in y_{t-1} (which would make it of size m), we should just get all the centroids at t-2 instead (all_curr_states[x-2]) which will be of size k, and m =/= n =/= k (at least we can assume so because it is not necessarily meant to be)\n\nnow these shapes are obviously incompatible (n, m, k), but the idea is to compute the smoothness cost, taken into account each triple combinations in the size m, n and k centroids arrays .\n\nSo y_t can be expanded to shape (n, 1, 1), then y_{t-1} can be expanded to shape (1, m, 1) and y_{t-2}will be (1, 1, k), this way, every shape the arrays broadcast to one another to give us a cost array of size (n, m, k), then add this to the distance cost of shape (n, m, 1) and compute the minimum along the last 2 axes. Hope that explains it.\n\n**PS: This function does need some improvements and is not meant to be used blindly as is, its just for demonstration. If you or anyone else have suggestions, do share them, I too am looking for an even better version or alternative.**",
    "3378368": "This is a genuinely insightful contribution thank you for taking the time to explain both the idea and the reasoning behind it so clearly.",
    "3378376": "You're welcome.",
    "3378480": "Thank you for the explanation!!!\n I will use it as a reference!",
    "3379734": "Thanks bro this will help me a lot"
  },
  "source": "meta"
}