{
  "id": 520479,
  "title": "Some parts of the metric didn't make sense to me",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/520479",
  "author_name": "",
  "post_date": "2024-07-16T07:53:51.892078500Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I was making a baseline with target means, playing with the predictions and metric but some parts didn't make sense.</p>\n<p>Predictions are stacked on Normal/Mild, Moderate and Severe columns and log losses are calculated on each condition group, but probabilities are not normalized in a way that their sum would be 1. It was like that in last years competition but not here. Is it intentional?</p>\n<p>I have doubts about any_severe_spinal implementation too. Labels, predictions and weights are max aggregated in study id_groups for spinal canal stenosis rows. Is it really supposed to be for only spinal canal stenosis rows or was that a mistake?</p>\n<p>If a study_id has one Moderate and four Normal/Mild labels for spinal canal stenosis, then the max weight becomes 2. Shouldn't the weights for any_severe_spinal be either 1 or 4 since the label name any_severe_spinal implies a binary prediction derived from other predictions. I don't think including Moderate labels on any_severe_spinal makes sense.</p>\n<p>What is any_severe_scalar? I couldn't find its value.</p>",
  "messages": [
    {
      "id": "2923910",
      "postDate": "07/16/2024 07:53:51",
      "content": "<p>I was making a baseline with target means, playing with the predictions and metric but some parts didn't make sense.</p>\n<p>Predictions are stacked on Normal/Mild, Moderate and Severe columns and log losses are calculated on each condition group, but probabilities are not normalized in a way that their sum would be 1. It was like that in last years competition but not here. Is it intentional?</p>\n<p>I have doubts about any_severe_spinal implementation too. Labels, predictions and weights are max aggregated in study id_groups for spinal canal stenosis rows. Is it really supposed to be for only spinal canal stenosis rows or was that a mistake?</p>\n<p>If a study_id has one Moderate and four Normal/Mild labels for spinal canal stenosis, then the max weight becomes 2. Shouldn't the weights for any_severe_spinal be either 1 or 4 since the label name any_severe_spinal implies a binary prediction derived from other predictions. I don't think including Moderate labels on any_severe_spinal makes sense.</p>\n<p>What is any_severe_scalar? I couldn't find its value.</p>",
      "rawMarkdown": "I was making a baseline with target means, playing with the predictions and metric but some parts didn't make sense.\n\nPredictions are stacked on Normal/Mild, Moderate and Severe columns and log losses are calculated on each condition group, but probabilities are not normalized in a way that their sum would be 1. It was like that in last years competition but not here. Is it intentional?\n\nI have doubts about any_severe_spinal implementation too. Labels, predictions and weights are max aggregated in study id_groups for spinal canal stenosis rows. Is it really supposed to be for only spinal canal stenosis rows or was that a mistake?\n\nIf a study_id has one Moderate and four Normal/Mild labels for spinal canal stenosis, then the max weight becomes 2. Shouldn't the weights for any_severe_spinal be either 1 or 4 since the label name any_severe_spinal implies a binary prediction derived from other predictions. I don't think including Moderate labels on any_severe_spinal makes sense.\n\nWhat is any_severe_scalar? I couldn't find its value.",
      "votes": null
    },
    {
      "id": "2924012",
      "postDate": "07/16/2024 08:52:35",
      "content": "<blockquote>\n  <p>What is any_severe_scalar? I couldn't find its value.</p>\n</blockquote>\n<p>According to the <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/overview/evaluation\" target=\"_blank\">Evaluation</a> tab:</p>\n<blockquote>\n  <p>For this competition, the any_severe_scalar has been set to 1.0.</p>\n</blockquote>\n<p>Edit:<br>\nMaybe this (formerly pinned) discussions would help you:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/510363\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/510363</a></li>\n</ul>",
      "rawMarkdown": "> What is any_severe_scalar? I couldn't find its value.\n\nAccording to the [Evaluation](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/overview/evaluation) tab:\n> For this competition, the any_severe_scalar has been set to 1.0.\n\nEdit:\nMaybe this (formerly pinned) discussions would help you:\n- https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522\n- https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/510363",
      "votes": null
    },
    {
      "id": "2924278",
      "postDate": "07/16/2024 12:36:58",
      "content": "<p>Thanks for the reply. Do you know if the probabilities are normalized or not?</p>",
      "rawMarkdown": "Thanks for the reply. Do you know if the probabilities are normalized or not?",
      "votes": null
    },
    {
      "id": "2924964",
      "postDate": "07/16/2024 20:27:21",
      "content": "<p>You can take a look at the metric code: <a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549\" target=\"_blank\">https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549</a></p>\n<p>The labels and predictions are directly inputted into <code>sklearn.metrics.log_loss</code>. Not sure if normalization occurs inside this function. If not, then no.</p>",
      "rawMarkdown": "You can take a look at the metric code: https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549\n\nThe labels and predictions are directly inputted into `sklearn.metrics.log_loss`. Not sure if normalization occurs inside this function. If not, then no.",
      "votes": null
    },
    {
      "id": "2929141",
      "postDate": "07/19/2024 21:11:03",
      "content": "<p>I believe <code>sklearn.metrics.log_loss</code> does normalize the probabilities to sum to 1 before calculating the loss. You can try this code snippet to confirm:</p>\n<pre><code>import numpy as np\nimport torch\n\n sklearn import log_loss\n\n\ndef torch_log_loss(, t):\n    loss = -torch.(t.(), p.()).()\n    return loss.()\n\n\ndef (x, y, eps=e-):\n    return np.(x - y) &lt; eps\n\n\n# No normalization\ny = torch.((, ))\ny[...] = \ny = torch.(y)\np = torch.((, ))\n\nsk_loss = (y.(), p.())\npt_loss = (p, y).()\n\n((sk_loss, pt_loss))\n\n# After normalization\np = p / p.().()\n\nsk_loss = (y.(), p.())\npt_loss = (p, y).()\n\n((sk_loss, pt_loss))\n</code></pre>\n<p>The values are only equal after you normalize the probabilities to sum to 1 (since <code>torch_log_loss</code> does not apply any normalization within the function).</p>",
      "rawMarkdown": "I believe `sklearn.metrics.log_loss` does normalize the probabilities to sum to 1 before calculating the loss. You can try this code snippet to confirm:\n\n```\nimport numpy as np\nimport torch\n\nfrom sklearn.metrics import log_loss\n\n\ndef torch_log_loss(p, t):\n\tloss = -torch.xlogy(t.float(), p.float()).sum(1)\n\treturn loss.mean()\n\n\ndef check_equal(x, y, eps=1e-6):\n\treturn np.abs(x - y) < eps\n\n\n# No normalization\ny = torch.empty((1000, 3))\ny[...] = 0.5\ny = torch.bernoulli(y)\np = torch.rand((1000, 3))\n\nsk_loss = log_loss(y.numpy(), p.numpy())\npt_loss = torch_log_loss(p, y).item()\n\nprint(check_equal(sk_loss, pt_loss))\n\n# After normalization\np = p / p.sum(1).unsqueeze(1)\n\nsk_loss = log_loss(y.numpy(), p.numpy())\npt_loss = torch_log_loss(p, y).item()\n\nprint(check_equal(sk_loss, pt_loss))\n```\n\nThe values are only equal after you normalize the probabilities to sum to 1 (since `torch_log_loss` does not apply any normalization within the function).",
      "votes": null
    },
    {
      "id": "2929188",
      "postDate": "07/19/2024 22:15:59",
      "content": "<p>Our metric is in a notebook that still uses v1.2.2, so it normalizes inputs. However, <code>sklearn</code> changed the normalization behavior in v1.3 so if you run a test off Kaggle you'll probably see the newer behavior: <a href=\"https://scikit-learn.org/stable/whats_new/v1.3.html\" target=\"_blank\">https://scikit-learn.org/stable/whats_new/v1.3.html</a>. </p>",
      "rawMarkdown": "Our metric is in a notebook that still uses v1.2.2, so it normalizes inputs. However, `sklearn` changed the normalization behavior in v1.3 so if you run a test off Kaggle you'll probably see the newer behavior: https://scikit-learn.org/stable/whats_new/v1.3.html.",
      "votes": null
    },
    {
      "id": "2929458",
      "postDate": "07/20/2024 06:21:28",
      "content": "<p>Thanks, I was looking for this answer. <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
      "rawMarkdown": "Thanks, I was looking for this answer. @vaillant @sohier",
      "votes": null
    },
    {
      "id": "2930546",
      "postDate": "07/21/2024 06:02:50",
      "content": "<p>I have a thought about any_severe_spinal weights implementation. Maybe for your example with moderate label, we have higher weight because it's harder to distinguish between moderate and severe case?</p>",
      "rawMarkdown": "I have a thought about any_severe_spinal weights implementation. Maybe for your example with moderate label, we have higher weight because it's harder to distinguish between moderate and severe case?",
      "votes": null
    },
    {
      "id": "2931848",
      "postDate": "07/22/2024 11:51:54",
      "content": "<p>I was using v1.3.2 when testing the code above, so it appears normalization still occurs.</p>",
      "rawMarkdown": "I was using v1.3.2 when testing the code above, so it appears normalization still occurs.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2924012,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "07/16/2024 08:52:35",
      "content": "<blockquote>\n  <p>What is any_severe_scalar? I couldn't find its value.</p>\n</blockquote>\n<p>According to the <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/overview/evaluation\" target=\"_blank\">Evaluation</a> tab:</p>\n<blockquote>\n  <p>For this competition, the any_severe_scalar has been set to 1.0.</p>\n</blockquote>\n<p>Edit:<br>\nMaybe this (formerly pinned) discussions would help you:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/510363\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/510363</a></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2924278,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "07/16/2024 12:36:58",
          "content": "<p>Thanks for the reply. Do you know if the probabilities are normalized or not?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2924964,
              "author_name": "coderrkj",
              "author_url": "",
              "post_date": "07/16/2024 20:27:21",
              "content": "<p>You can take a look at the metric code: <a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549\" target=\"_blank\">https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549</a></p>\n<p>The labels and predictions are directly inputted into <code>sklearn.metrics.log_loss</code>. Not sure if normalization occurs inside this function. If not, then no.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2929141,
                  "author_name": "vaillant",
                  "author_url": "",
                  "post_date": "07/19/2024 21:11:03",
                  "content": "<p>I believe <code>sklearn.metrics.log_loss</code> does normalize the probabilities to sum to 1 before calculating the loss. You can try this code snippet to confirm:</p>\n<pre><code>import numpy as np\nimport torch\n\n sklearn import log_loss\n\n\ndef torch_log_loss(, t):\n    loss = -torch.(t.(), p.()).()\n    return loss.()\n\n\ndef (x, y, eps=e-):\n    return np.(x - y) &lt; eps\n\n\n# No normalization\ny = torch.((, ))\ny[...] = \ny = torch.(y)\np = torch.((, ))\n\nsk_loss = (y.(), p.())\npt_loss = (p, y).()\n\n((sk_loss, pt_loss))\n\n# After normalization\np = p / p.().()\n\nsk_loss = (y.(), p.())\npt_loss = (p, y).()\n\n((sk_loss, pt_loss))\n</code></pre>\n<p>The values are only equal after you normalize the probabilities to sum to 1 (since <code>torch_log_loss</code> does not apply any normalization within the function).</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2929188,
                      "author_name": "sohier",
                      "author_url": "",
                      "post_date": "07/19/2024 22:15:59",
                      "content": "<p>Our metric is in a notebook that still uses v1.2.2, so it normalizes inputs. However, <code>sklearn</code> changed the normalization behavior in v1.3 so if you run a test off Kaggle you'll probably see the newer behavior: <a href=\"https://scikit-learn.org/stable/whats_new/v1.3.html\" target=\"_blank\">https://scikit-learn.org/stable/whats_new/v1.3.html</a>. </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2929458,
                          "author_name": "gunesevitan",
                          "author_url": "",
                          "post_date": "07/20/2024 06:21:28",
                          "content": "<p>Thanks, I was looking for this answer. <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2931848,
                              "author_name": "vaillant",
                              "author_url": "",
                              "post_date": "07/22/2024 11:51:54",
                              "content": "<p>I was using v1.3.2 when testing the code above, so it appears normalization still occurs.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2930546,
      "author_name": "glipko",
      "author_url": "",
      "post_date": "07/21/2024 06:02:50",
      "content": "<p>I have a thought about any_severe_spinal weights implementation. Maybe for your example with moderate label, we have higher weight because it's harder to distinguish between moderate and severe case?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2923910": "I was making a baseline with target means, playing with the predictions and metric but some parts didn't make sense.\n\nPredictions are stacked on Normal/Mild, Moderate and Severe columns and log losses are calculated on each condition group, but probabilities are not normalized in a way that their sum would be 1. It was like that in last years competition but not here. Is it intentional?\n\nI have doubts about any_severe_spinal implementation too. Labels, predictions and weights are max aggregated in study id_groups for spinal canal stenosis rows. Is it really supposed to be for only spinal canal stenosis rows or was that a mistake?\n\nIf a study_id has one Moderate and four Normal/Mild labels for spinal canal stenosis, then the max weight becomes 2. Shouldn't the weights for any_severe_spinal be either 1 or 4 since the label name any_severe_spinal implies a binary prediction derived from other predictions. I don't think including Moderate labels on any_severe_spinal makes sense.\n\nWhat is any_severe_scalar? I couldn't find its value.",
    "2924012": "> What is any_severe_scalar? I couldn't find its value.\n\nAccording to the [Evaluation](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/overview/evaluation) tab:\n> For this competition, the any_severe_scalar has been set to 1.0.\n\nEdit:\nMaybe this (formerly pinned) discussions would help you:\n- https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522\n- https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/510363",
    "2924278": "Thanks for the reply. Do you know if the probabilities are normalized or not?",
    "2924964": "You can take a look at the metric code: https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549\n\nThe labels and predictions are directly inputted into `sklearn.metrics.log_loss`. Not sure if normalization occurs inside this function. If not, then no.",
    "2929141": "I believe `sklearn.metrics.log_loss` does normalize the probabilities to sum to 1 before calculating the loss. You can try this code snippet to confirm:\n\n```\nimport numpy as np\nimport torch\n\nfrom sklearn.metrics import log_loss\n\n\ndef torch_log_loss(p, t):\n\tloss = -torch.xlogy(t.float(), p.float()).sum(1)\n\treturn loss.mean()\n\n\ndef check_equal(x, y, eps=1e-6):\n\treturn np.abs(x - y) < eps\n\n\n# No normalization\ny = torch.empty((1000, 3))\ny[...] = 0.5\ny = torch.bernoulli(y)\np = torch.rand((1000, 3))\n\nsk_loss = log_loss(y.numpy(), p.numpy())\npt_loss = torch_log_loss(p, y).item()\n\nprint(check_equal(sk_loss, pt_loss))\n\n# After normalization\np = p / p.sum(1).unsqueeze(1)\n\nsk_loss = log_loss(y.numpy(), p.numpy())\npt_loss = torch_log_loss(p, y).item()\n\nprint(check_equal(sk_loss, pt_loss))\n```\n\nThe values are only equal after you normalize the probabilities to sum to 1 (since `torch_log_loss` does not apply any normalization within the function).",
    "2929188": "Our metric is in a notebook that still uses v1.2.2, so it normalizes inputs. However, `sklearn` changed the normalization behavior in v1.3 so if you run a test off Kaggle you'll probably see the newer behavior: https://scikit-learn.org/stable/whats_new/v1.3.html.",
    "2929458": "Thanks, I was looking for this answer. @vaillant @sohier",
    "2930546": "I have a thought about any_severe_spinal weights implementation. Maybe for your example with moderate label, we have higher weight because it's harder to distinguish between moderate and severe case?",
    "2931848": "I was using v1.3.2 when testing the code above, so it appears normalization still occurs."
  },
  "source": "meta"
}