{
  "id": 663901,
  "title": "Critical Metric Vulnerability: One artifact can overwhelm the remaining 999 samples",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/663901",
  "author_name": "",
  "post_date": "2025-12-21T11:11:21.994374500Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Dear Organizers,</p>\n<p>I am writing to report a critical mathematical vulnerability in the current evaluation metric (based on<code>Mean Linear SNR</code>) that severely compromises the competition's goal of creating a robust model.</p>\n<p>While one might assume that having ~1,000 test samples would average out outliers, the exponential nature of the Linear SNR ( signal^2/noise^2 ) makes the sample size irrelevant in the face of high-amplitude artifacts.</p>\n<p><strong>0. Current Metrics Computation</strong></p>\n<ol>\n<li>calculate linear SNR (signal power / noise power) per image</li>\n<li>take mean of 1, and convert to decibel</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe04f4fee00e2fa176d27ae1adee5505a%2FScreenshot%202025-12-21%20at%2020.09.13.png?generation=1766315369548557&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric/\" target=\"_blank\">https://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric/</a></p>\n<p><strong>1. The \"1 vs 100\" Effort Imbalance</strong></p>\n<ul>\n<li><strong>Normal Sample (20 dB):</strong> Linear Ratio = 100 x 999</li>\n<li><strong>Good/Artifact Sample (40 dB):</strong> Linear Ratio = 10,000 x 1</li>\n</ul>\n<p>Even if a single outlier doesn't ruin the whole score immediately, the <strong>incentive structure</strong> is broken.\nMathematically, improving <strong>ONE single sample</strong> to 40 dB ( ratio) yields the same leaderboard boost (+0.4 dB) as improving <strong>ONE HUNDRED samples</strong> by 3 dB (doubling their accuracy).</p>\n<ul>\n<li><strong>Path A (Hacking):</strong> Overfit to 1 artifact image to get 40dB  -&gt; <strong>+0.4 dB gain</strong></li>\n</ul>\n<pre><code>10 * np.log10((10_000 + 100 * 999) / 1_000)\n# Result: np.float64(20.409976924234904)\n</code></pre>\n<ul>\n<li><strong>Path B (Research):</strong> Significantly improve model architecture to handle 100 different patients better.  -&gt; <strong>+0.4 dB gain</strong></li>\n</ul>\n<pre><code>10 * np.log10((200 * 100 + 100 * 900) / 1_000)\n# Result: np.float64(20.41392685158225)\n</code></pre>\n<p>This forces participants to choose Path A (chasing artifacts) over Path B (building a better model), which contradicts the competition's goal.</p>\n<p><strong>2. The \"Instant Win\" Scenario (The 60 dB Exploit)</strong></p>\n<ul>\n<li><strong>Normal Sample (20 dB):</strong> Linear Ratio = 100 x 999</li>\n<li><strong>Good/Artifact Sample (60 dB):</strong> Linear Ratio = 1,000,000 x1</li>\n</ul>\n<p>It's difficult to achive 60dB, but if one can get 60 dB for only 1 sample, this participant can score 30dB, even if he/she gets 0 dB for the rest of the samples.</p>\n<pre><code>10 * np.log10((1_000_000 + 1 * 999) / 1_000)\n# Result: ~30.004 dB -&gt; approx 30 dB\n</code></pre>\n<p><strong>3. The Solution: Global SNR</strong>\nTo fix this vulnerability and reward true robustness, I strongly propose changing the metric to <strong>Global SNR</strong>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F9b6af7df1dd099ec9002df8e9769b817%2FScreenshot%202025-12-21%20at%2019.52.34.png?generation=1766314370594324&amp;alt=media\" alt=\"\"></p>\n<p>This method sums the signal powers and noise powers first, treating the dataset as one continuous signal. It prevents any single sample from dominating the score and ensures that the winner is actually the most robust solution for the entire dataset.</p>\n<p>Best regards,</p>\n<p><a href=\"https://www.kaggle.com/gdclifford\" target=\"_blank\">@gdclifford</a>\n<a href=\"https://www.kaggle.com/r2241272\" target=\"_blank\">@r2241272</a></p>",
  "messages": [
    {
      "id": "3379984",
      "postDate": "12/21/2025 11:11:21",
      "content": "<p>Dear Organizers,</p>\n<p>I am writing to report a critical mathematical vulnerability in the current evaluation metric (based on<code>Mean Linear SNR</code>) that severely compromises the competition's goal of creating a robust model.</p>\n<p>While one might assume that having ~1,000 test samples would average out outliers, the exponential nature of the Linear SNR ( signal^2/noise^2 ) makes the sample size irrelevant in the face of high-amplitude artifacts.</p>\n<p><strong>0. Current Metrics Computation</strong></p>\n<ol>\n<li>calculate linear SNR (signal power / noise power) per image</li>\n<li>take mean of 1, and convert to decibel</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe04f4fee00e2fa176d27ae1adee5505a%2FScreenshot%202025-12-21%20at%2020.09.13.png?generation=1766315369548557&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric/\" target=\"_blank\">https://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric/</a></p>\n<p><strong>1. The \"1 vs 100\" Effort Imbalance</strong></p>\n<ul>\n<li><strong>Normal Sample (20 dB):</strong> Linear Ratio = 100 x 999</li>\n<li><strong>Good/Artifact Sample (40 dB):</strong> Linear Ratio = 10,000 x 1</li>\n</ul>\n<p>Even if a single outlier doesn't ruin the whole score immediately, the <strong>incentive structure</strong> is broken.\nMathematically, improving <strong>ONE single sample</strong> to 40 dB ( ratio) yields the same leaderboard boost (+0.4 dB) as improving <strong>ONE HUNDRED samples</strong> by 3 dB (doubling their accuracy).</p>\n<ul>\n<li><strong>Path A (Hacking):</strong> Overfit to 1 artifact image to get 40dB  -&gt; <strong>+0.4 dB gain</strong></li>\n</ul>\n<pre><code>10 * np.log10((10_000 + 100 * 999) / 1_000)\n# Result: np.float64(20.409976924234904)\n</code></pre>\n<ul>\n<li><strong>Path B (Research):</strong> Significantly improve model architecture to handle 100 different patients better.  -&gt; <strong>+0.4 dB gain</strong></li>\n</ul>\n<pre><code>10 * np.log10((200 * 100 + 100 * 900) / 1_000)\n# Result: np.float64(20.41392685158225)\n</code></pre>\n<p>This forces participants to choose Path A (chasing artifacts) over Path B (building a better model), which contradicts the competition's goal.</p>\n<p><strong>2. The \"Instant Win\" Scenario (The 60 dB Exploit)</strong></p>\n<ul>\n<li><strong>Normal Sample (20 dB):</strong> Linear Ratio = 100 x 999</li>\n<li><strong>Good/Artifact Sample (60 dB):</strong> Linear Ratio = 1,000,000 x1</li>\n</ul>\n<p>It's difficult to achive 60dB, but if one can get 60 dB for only 1 sample, this participant can score 30dB, even if he/she gets 0 dB for the rest of the samples.</p>\n<pre><code>10 * np.log10((1_000_000 + 1 * 999) / 1_000)\n# Result: ~30.004 dB -&gt; approx 30 dB\n</code></pre>\n<p><strong>3. The Solution: Global SNR</strong>\nTo fix this vulnerability and reward true robustness, I strongly propose changing the metric to <strong>Global SNR</strong>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F9b6af7df1dd099ec9002df8e9769b817%2FScreenshot%202025-12-21%20at%2019.52.34.png?generation=1766314370594324&amp;alt=media\" alt=\"\"></p>\n<p>This method sums the signal powers and noise powers first, treating the dataset as one continuous signal. It prevents any single sample from dominating the score and ensures that the winner is actually the most robust solution for the entire dataset.</p>\n<p>Best regards,</p>\n<p><a href=\"https://www.kaggle.com/gdclifford\" target=\"_blank\">@gdclifford</a>\n<a href=\"https://www.kaggle.com/r2241272\" target=\"_blank\">@r2241272</a></p>",
      "rawMarkdown": "Dear Organizers,\n\nI am writing to report a critical mathematical vulnerability in the current evaluation metric (based on`Mean Linear SNR`) that severely compromises the competition's goal of creating a robust model.\n\nWhile one might assume that having ~1,000 test samples would average out outliers, the exponential nature of the Linear SNR ( signal^2/noise^2 ) makes the sample size irrelevant in the face of high-amplitude artifacts.\n\n**0. Current Metrics Computation**\n\n1. calculate linear SNR (signal power / noise power) per image\n2. take mean of 1, and convert to decibel\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe04f4fee00e2fa176d27ae1adee5505a%2FScreenshot%202025-12-21%20at%2020.09.13.png?generation=1766315369548557&alt=media)\n\nhttps://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric/\n\n\n**1. The \"1 vs 100\" Effort Imbalance**\n\n* **Normal Sample (20 dB):** Linear Ratio = 100 x 999\n* **Good/Artifact Sample (40 dB):** Linear Ratio = 10,000 x 1\n\n\nEven if a single outlier doesn't ruin the whole score immediately, the **incentive structure** is broken.\nMathematically, improving **ONE single sample** to 40 dB ( ratio) yields the same leaderboard boost (+0.4 dB) as improving **ONE HUNDRED samples** by 3 dB (doubling their accuracy).\n\n* **Path A (Hacking):** Overfit to 1 artifact image to get 40dB  -> **+0.4 dB gain**\n\n```python\n10 * np.log10((10_000 + 100 * 999) / 1_000)\n# Result: np.float64(20.409976924234904)\n```\n\n* **Path B (Research):** Significantly improve model architecture to handle 100 different patients better.  -> **+0.4 dB gain**\n\n```python\n10 * np.log10((200 * 100 + 100 * 900) / 1_000)\n# Result: np.float64(20.41392685158225)\n```\n\nThis forces participants to choose Path A (chasing artifacts) over Path B (building a better model), which contradicts the competition's goal.\n\n**2. The \"Instant Win\" Scenario (The 60 dB Exploit)**\n\n* **Normal Sample (20 dB):** Linear Ratio = 100 x 999\n* **Good/Artifact Sample (60 dB):** Linear Ratio = 1,000,000 x1\n\nIt's difficult to achive 60dB, but if one can get 60 dB for only 1 sample, this participant can score 30dB, even if he/she gets 0 dB for the rest of the samples.\n\n```python\n10 * np.log10((1_000_000 + 1 * 999) / 1_000)\n# Result: ~30.004 dB -> approx 30 dB\n```\n\n**3. The Solution: Global SNR**\nTo fix this vulnerability and reward true robustness, I strongly propose changing the metric to **Global SNR**:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F9b6af7df1dd099ec9002df8e9769b817%2FScreenshot%202025-12-21%20at%2019.52.34.png?generation=1766314370594324&alt=media)\n\nThis method sums the signal powers and noise powers first, treating the dataset as one continuous signal. It prevents any single sample from dominating the score and ensures that the winner is actually the most robust solution for the entire dataset.\n\nBest regards,\n\n@gdclifford\n@r2241272",
      "votes": null
    },
    {
      "id": "3380026",
      "postDate": "12/21/2025 13:25:41",
      "content": "<p>Thank you for bringing this up and for the nice discussion. We discussed and simulated several different ways to define the score when the challenge was designed, including your proposed score. Due to the multi-lead and multi-subject nature of the problem, and the nonlinearity of the SNR, there is no single way to define the score. We decided on the current score for the following reason: the current score calculates the SNR per image, then averages the SNR across all images and converts it into dB; your approach, would be equivalent to stitching all images into a single image and calculating the SNR across that large image. However, cross-image behavior is practically more important for the challenge and its generalizability in clinical settings. Each ECG comes separately from an independent subject and has its own imaging artifacts. For example, rotation and tilts, monitor artifacts, wrinkles, reflections, dirt, etc. are different issues, and we do not expect all algorithms to work equally well across all of these artifacts.</p>\n<p>In practice, as an operator of an ECG digitization system, you would rather identify images on which the model does not work well and have them overread by a human or manually fixed by hand tracing, rather than use a model that performs mediocrely across all subjects and ECGs. Of course, a model that does a good job on every single image will also perform well on average.</p>\n<p>Another reason for choosing the current score is that there are cases in which we know no model will perform well, including situations where the tracings exceed the page limit due to excessive baseline wander or device failure. While not very common, this does happen in practice, and no digitization software can resolve it (see <a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/613040\" target=\"_blank\">this discussion</a>). The current score is less impacted by isolated problematic cases like this.</p>\n<p>In the end, there is no single correct way to define the metric, and I agree that your method has advantages as well. However, the challenge score will remain as is for now, and we can leave offline tests of susceptibility to corner cases for after the challenge.</p>",
      "rawMarkdown": "Thank you for bringing this up and for the nice discussion. We discussed and simulated several different ways to define the score when the challenge was designed, including your proposed score. Due to the multi-lead and multi-subject nature of the problem, and the nonlinearity of the SNR, there is no single way to define the score. We decided on the current score for the following reason: the current score calculates the SNR per image, then averages the SNR across all images and converts it into dB; your approach, would be equivalent to stitching all images into a single image and calculating the SNR across that large image. However, cross-image behavior is practically more important for the challenge and its generalizability in clinical settings. Each ECG comes separately from an independent subject and has its own imaging artifacts. For example, rotation and tilts, monitor artifacts, wrinkles, reflections, dirt, etc. are different issues, and we do not expect all algorithms to work equally well across all of these artifacts.\n\nIn practice, as an operator of an ECG digitization system, you would rather identify images on which the model does not work well and have them overread by a human or manually fixed by hand tracing, rather than use a model that performs mediocrely across all subjects and ECGs. Of course, a model that does a good job on every single image will also perform well on average.\n\nAnother reason for choosing the current score is that there are cases in which we know no model will perform well, including situations where the tracings exceed the page limit due to excessive baseline wander or device failure. While not very common, this does happen in practice, and no digitization software can resolve it (see [this discussion](https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/613040)). The current score is less impacted by isolated problematic cases like this.\n\nIn the end, there is no single correct way to define the metric, and I agree that your method has advantages as well. However, the challenge score will remain as is for now, and we can leave offline tests of susceptibility to corner cases for after the challenge.",
      "votes": null
    },
    {
      "id": "3380096",
      "postDate": "12/21/2025 16:19:36",
      "content": "<p><a href=\"https://www.kaggle.com/r2241272\" target=\"_blank\">@r2241272</a> I see. Thank you for the clarification.</p>",
      "rawMarkdown": "r2241272 I see. Thank you for the clarification.",
      "votes": null
    },
    {
      "id": "3382796",
      "postDate": "12/28/2025 15:36:14",
      "content": "<p>The metric actually has a much bigger problem.\nIf you shift one QRS segment by 1 sample left and the other one 1 sample right (completely reasonable, does not change interpretation at all), you get a big error.\nIf you remove the P-wave (that small notch before the QRS), you get a small error.</p>",
      "rawMarkdown": "The metric actually has a much bigger problem.\nIf you shift one QRS segment by 1 sample left and the other one 1 sample right (completely reasonable, does not change interpretation at all), you get a big error.\nIf you remove the P-wave (that small notch before the QRS), you get a small error.",
      "votes": null
    },
    {
      "id": "3383168",
      "postDate": "12/29/2025 14:58:17",
      "content": "<p>The challenge score already compensates for minor horizontal and vertical shifts. However, it's true that different segments of the ECG do not carry the same diagnostic value. For example, one could use a weighted SNR that prioritizes segments with higher diagnostic importance (or like what George Moody humbly discussed <a href=\"https://physionet.org/physiotools/wag/nst-1.htm\" target=\"_blank\">here</a> about the QRS complex SNR); but that would require R-peak detection, ECG segmentation, fiducial point extraction and other processing steps. This, in turn, would entangle the digitization task with the underlying ECG pathology, signal quality and the performance of the delineation algorithms. That's why, in the <a href=\"https://moody-challenge.physionet.org/2024/\" target=\"_blank\">PhysioNet Challenge 2024</a>, we defined two complementary but independent tasks: digitization and diagnosis. Teams could choose either to first digitize the ECG and then perform diagnosis, or to go directly from an ECG image to diagnosis.</p>\n<p>Since this Kaggle challenge focuses only on the first task, the scoring metric needed to be objective and largely independent of the dataset and underlying pathologies. That's why we decided to use the current objective metric. Note also that at high SNRs, even subtle details such as the P and Q waves contribute to the score. So, when an algorithm extracts an ECG with an SNR above 20–30dB, it is also performing quite well on low-amplitude segments such as the P and Q waves. Nonetheless, such sensitivity analysis is an open problem that we plan to investigate after the challenge.</p>",
      "rawMarkdown": "The challenge score already compensates for minor horizontal and vertical shifts. However, it's true that different segments of the ECG do not carry the same diagnostic value. For example, one could use a weighted SNR that prioritizes segments with higher diagnostic importance (or like what George Moody humbly discussed [here](https://physionet.org/physiotools/wag/nst-1.htm) about the QRS complex SNR); but that would require R-peak detection, ECG segmentation, fiducial point extraction and other processing steps. This, in turn, would entangle the digitization task with the underlying ECG pathology, signal quality and the performance of the delineation algorithms. That's why, in the [PhysioNet Challenge 2024](https://moody-challenge.physionet.org/2024/), we defined two complementary but independent tasks: digitization and diagnosis. Teams could choose either to first digitize the ECG and then perform diagnosis, or to go directly from an ECG image to diagnosis.\n\nSince this Kaggle challenge focuses only on the first task, the scoring metric needed to be objective and largely independent of the dataset and underlying pathologies. That's why we decided to use the current objective metric. Note also that at high SNRs, even subtle details such as the P and Q waves contribute to the score. So, when an algorithm extracts an ECG with an SNR above 20–30dB, it is also performing quite well on low-amplitude segments such as the P and Q waves. Nonetheless, such sensitivity analysis is an open problem that we plan to investigate after the challenge.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3380026,
      "author_name": "r2241272",
      "author_url": "",
      "post_date": "12/21/2025 13:25:41",
      "content": "<p>Thank you for bringing this up and for the nice discussion. We discussed and simulated several different ways to define the score when the challenge was designed, including your proposed score. Due to the multi-lead and multi-subject nature of the problem, and the nonlinearity of the SNR, there is no single way to define the score. We decided on the current score for the following reason: the current score calculates the SNR per image, then averages the SNR across all images and converts it into dB; your approach, would be equivalent to stitching all images into a single image and calculating the SNR across that large image. However, cross-image behavior is practically more important for the challenge and its generalizability in clinical settings. Each ECG comes separately from an independent subject and has its own imaging artifacts. For example, rotation and tilts, monitor artifacts, wrinkles, reflections, dirt, etc. are different issues, and we do not expect all algorithms to work equally well across all of these artifacts.</p>\n<p>In practice, as an operator of an ECG digitization system, you would rather identify images on which the model does not work well and have them overread by a human or manually fixed by hand tracing, rather than use a model that performs mediocrely across all subjects and ECGs. Of course, a model that does a good job on every single image will also perform well on average.</p>\n<p>Another reason for choosing the current score is that there are cases in which we know no model will perform well, including situations where the tracings exceed the page limit due to excessive baseline wander or device failure. While not very common, this does happen in practice, and no digitization software can resolve it (see <a href=\"https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/613040\" target=\"_blank\">this discussion</a>). The current score is less impacted by isolated problematic cases like this.</p>\n<p>In the end, there is no single correct way to define the metric, and I agree that your method has advantages as well. However, the challenge score will remain as is for now, and we can leave offline tests of susceptibility to corner cases for after the challenge.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3380096,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "12/21/2025 16:19:36",
          "content": "<p><a href=\"https://www.kaggle.com/r2241272\" target=\"_blank\">@r2241272</a> I see. Thank you for the clarification.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3382796,
      "author_name": "usamec",
      "author_url": "",
      "post_date": "12/28/2025 15:36:14",
      "content": "<p>The metric actually has a much bigger problem.\nIf you shift one QRS segment by 1 sample left and the other one 1 sample right (completely reasonable, does not change interpretation at all), you get a big error.\nIf you remove the P-wave (that small notch before the QRS), you get a small error.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3383168,
          "author_name": "r2241272",
          "author_url": "",
          "post_date": "12/29/2025 14:58:17",
          "content": "<p>The challenge score already compensates for minor horizontal and vertical shifts. However, it's true that different segments of the ECG do not carry the same diagnostic value. For example, one could use a weighted SNR that prioritizes segments with higher diagnostic importance (or like what George Moody humbly discussed <a href=\"https://physionet.org/physiotools/wag/nst-1.htm\" target=\"_blank\">here</a> about the QRS complex SNR); but that would require R-peak detection, ECG segmentation, fiducial point extraction and other processing steps. This, in turn, would entangle the digitization task with the underlying ECG pathology, signal quality and the performance of the delineation algorithms. That's why, in the <a href=\"https://moody-challenge.physionet.org/2024/\" target=\"_blank\">PhysioNet Challenge 2024</a>, we defined two complementary but independent tasks: digitization and diagnosis. Teams could choose either to first digitize the ECG and then perform diagnosis, or to go directly from an ECG image to diagnosis.</p>\n<p>Since this Kaggle challenge focuses only on the first task, the scoring metric needed to be objective and largely independent of the dataset and underlying pathologies. That's why we decided to use the current objective metric. Note also that at high SNRs, even subtle details such as the P and Q waves contribute to the score. So, when an algorithm extracts an ECG with an SNR above 20–30dB, it is also performing quite well on low-amplitude segments such as the P and Q waves. Nonetheless, such sensitivity analysis is an open problem that we plan to investigate after the challenge.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3379984": "Dear Organizers,\n\nI am writing to report a critical mathematical vulnerability in the current evaluation metric (based on`Mean Linear SNR`) that severely compromises the competition's goal of creating a robust model.\n\nWhile one might assume that having ~1,000 test samples would average out outliers, the exponential nature of the Linear SNR ( signal^2/noise^2 ) makes the sample size irrelevant in the face of high-amplitude artifacts.\n\n**0. Current Metrics Computation**\n\n1. calculate linear SNR (signal power / noise power) per image\n2. take mean of 1, and convert to decibel\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe04f4fee00e2fa176d27ae1adee5505a%2FScreenshot%202025-12-21%20at%2020.09.13.png?generation=1766315369548557&alt=media)\n\nhttps://www.kaggle.com/code/metric/physionet-ecg-signal-extraction-metric/\n\n\n**1. The \"1 vs 100\" Effort Imbalance**\n\n* **Normal Sample (20 dB):** Linear Ratio = 100 x 999\n* **Good/Artifact Sample (40 dB):** Linear Ratio = 10,000 x 1\n\n\nEven if a single outlier doesn't ruin the whole score immediately, the **incentive structure** is broken.\nMathematically, improving **ONE single sample** to 40 dB ( ratio) yields the same leaderboard boost (+0.4 dB) as improving **ONE HUNDRED samples** by 3 dB (doubling their accuracy).\n\n* **Path A (Hacking):** Overfit to 1 artifact image to get 40dB  -> **+0.4 dB gain**\n\n```python\n10 * np.log10((10_000 + 100 * 999) / 1_000)\n# Result: np.float64(20.409976924234904)\n```\n\n* **Path B (Research):** Significantly improve model architecture to handle 100 different patients better.  -> **+0.4 dB gain**\n\n```python\n10 * np.log10((200 * 100 + 100 * 900) / 1_000)\n# Result: np.float64(20.41392685158225)\n```\n\nThis forces participants to choose Path A (chasing artifacts) over Path B (building a better model), which contradicts the competition's goal.\n\n**2. The \"Instant Win\" Scenario (The 60 dB Exploit)**\n\n* **Normal Sample (20 dB):** Linear Ratio = 100 x 999\n* **Good/Artifact Sample (60 dB):** Linear Ratio = 1,000,000 x1\n\nIt's difficult to achive 60dB, but if one can get 60 dB for only 1 sample, this participant can score 30dB, even if he/she gets 0 dB for the rest of the samples.\n\n```python\n10 * np.log10((1_000_000 + 1 * 999) / 1_000)\n# Result: ~30.004 dB -> approx 30 dB\n```\n\n**3. The Solution: Global SNR**\nTo fix this vulnerability and reward true robustness, I strongly propose changing the metric to **Global SNR**:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F9b6af7df1dd099ec9002df8e9769b817%2FScreenshot%202025-12-21%20at%2019.52.34.png?generation=1766314370594324&alt=media)\n\nThis method sums the signal powers and noise powers first, treating the dataset as one continuous signal. It prevents any single sample from dominating the score and ensures that the winner is actually the most robust solution for the entire dataset.\n\nBest regards,\n\n@gdclifford\n@r2241272",
    "3380026": "Thank you for bringing this up and for the nice discussion. We discussed and simulated several different ways to define the score when the challenge was designed, including your proposed score. Due to the multi-lead and multi-subject nature of the problem, and the nonlinearity of the SNR, there is no single way to define the score. We decided on the current score for the following reason: the current score calculates the SNR per image, then averages the SNR across all images and converts it into dB; your approach, would be equivalent to stitching all images into a single image and calculating the SNR across that large image. However, cross-image behavior is practically more important for the challenge and its generalizability in clinical settings. Each ECG comes separately from an independent subject and has its own imaging artifacts. For example, rotation and tilts, monitor artifacts, wrinkles, reflections, dirt, etc. are different issues, and we do not expect all algorithms to work equally well across all of these artifacts.\n\nIn practice, as an operator of an ECG digitization system, you would rather identify images on which the model does not work well and have them overread by a human or manually fixed by hand tracing, rather than use a model that performs mediocrely across all subjects and ECGs. Of course, a model that does a good job on every single image will also perform well on average.\n\nAnother reason for choosing the current score is that there are cases in which we know no model will perform well, including situations where the tracings exceed the page limit due to excessive baseline wander or device failure. While not very common, this does happen in practice, and no digitization software can resolve it (see [this discussion](https://www.kaggle.com/competitions/physionet-ecg-image-digitization/discussion/613040)). The current score is less impacted by isolated problematic cases like this.\n\nIn the end, there is no single correct way to define the metric, and I agree that your method has advantages as well. However, the challenge score will remain as is for now, and we can leave offline tests of susceptibility to corner cases for after the challenge.",
    "3380096": "r2241272 I see. Thank you for the clarification.",
    "3382796": "The metric actually has a much bigger problem.\nIf you shift one QRS segment by 1 sample left and the other one 1 sample right (completely reasonable, does not change interpretation at all), you get a big error.\nIf you remove the P-wave (that small notch before the QRS), you get a small error.",
    "3383168": "The challenge score already compensates for minor horizontal and vertical shifts. However, it's true that different segments of the ECG do not carry the same diagnostic value. For example, one could use a weighted SNR that prioritizes segments with higher diagnostic importance (or like what George Moody humbly discussed [here](https://physionet.org/physiotools/wag/nst-1.htm) about the QRS complex SNR); but that would require R-peak detection, ECG segmentation, fiducial point extraction and other processing steps. This, in turn, would entangle the digitization task with the underlying ECG pathology, signal quality and the performance of the delineation algorithms. That's why, in the [PhysioNet Challenge 2024](https://moody-challenge.physionet.org/2024/), we defined two complementary but independent tasks: digitization and diagnosis. Teams could choose either to first digitize the ECG and then perform diagnosis, or to go directly from an ECG image to diagnosis.\n\nSince this Kaggle challenge focuses only on the first task, the scoring metric needed to be objective and largely independent of the dataset and underlying pathologies. That's why we decided to use the current objective metric. Note also that at high SNRs, even subtle details such as the P and Q waves contribute to the score. So, when an algorithm extracts an ECG with an SNR above 20–30dB, it is also performing quite well on low-amplitude segments such as the P and Q waves. Nonetheless, such sensitivity analysis is an open problem that we plan to investigate after the challenge."
  },
  "source": "meta"
}