{
  "id": 515357,
  "title": "Metric Behavior Complaint v3",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/515357",
  "author_name": "",
  "post_date": "2024-06-27T20:27:28.052490900Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi,</p>\n<h2>Intro</h2>\n<hr>\n<p>I've been experimenting with <code>any_severe_spinal</code> part of the metric for a while and come to the conclusion that there is a chance for <code>random walking</code> of the metric. Let's look at the current metric implementation for that part below:</p>\n<pre><code>any_severe_spinal_labels = pd.Series(\n    solution.loc[solution[] == ].groupby()[].())\nany_severe_spinal_weights = pd.Series(\n    solution.loc[solution[] == ].groupby()[].())\nany_severe_spinal_predictions = pd.Series(\n    submission.loc[submission[] == ].groupby()[].())\nany_severe_spinal_loss = log_loss(\n    y_true=any_severe_spinal_labels,\n    y_pred=any_severe_spinal_predictions,\n    sample_weight=any_severe_spinal_weights\n)\n</code></pre>\n<p>What happens is that by taking <code>.max()</code> we apply it for 5 levels (L1/L2, L2/L3, L3/L4, L4/L5, L5/S1) for a particular <code>study_id</code>. <br>\nWhat if we mispredict some other level as <code>severe</code> with high probability but completely fail to predict the actual level with <code>severe</code> pathology? </p>\n<h2>Real world example</h2>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F4dcfc6ae788d5c151a2637ea909916bb%2FScreenshot%20from%202024-06-27%2015-54-38.png?generation=1719518100343248&amp;alt=media\"></p>\n<p>What we see here is that we failed to predict the <code>severe</code> label for <code>L4/L5</code> level, but <strong>with current implementation it is not a big deal!</strong> - the metric gives us a second chance. How does it happen? Well, the current implementation actually takes <code>.max()</code> which in this case is <code>L3/L4</code> with value <code>~0.52</code>. So now, rather than having <code>true</code> logloss equals to <code>~2.104</code> we get nice <code>~0.6538</code>. </p>\n<h2>Potential ways to fix it:</h2>\n<hr>\n<ol>\n<li>Taking only <code>spinal</code> &amp; <code>severe</code> cases:</li>\n</ol>\n<pre><code>any_severe_spinal_labels = solution.loc[(solution[] == ) &amp; (solution[] == )]\nsevere_spinal_indices = any_severe_spinal_labels.index\nany_severe_spinal_weights = solution.loc[severe_spinal_indices, ]\nany_severe_spinal_predictions = submission.loc[severe_spinal_indices]\nany_severe_spinal_loss = log_loss(\n    y_true=any_severe_spinal_labels[target_levels],\n    y_pred=any_severe_spinal_predictions[target_levels],\n    sample_weight=any_severe_spinal_weights\n) \n</code></pre>\n<p>The problem with that is that we want to take <code>any_severe_spinal</code> observation, where <code>any</code> stands for \"does not matter which level\" - if I understand correctly the host intention.</p>\n<ol>\n<li>Taking the first <code>spinal</code> &amp; <code>severe</code> case per study_id:</li>\n</ol>\n<pre><code>any_severe_spinal_labels = solution.loc[(solution[] == ) &amp; (solution.severe == )]\nany_severe_spinal_labels = any_severe_spinal_labels.drop_duplicates(subset=)\nsevere_spinal_indices = any_severe_spinal_labels.index\nany_severe_spinal_weights = solution.loc[severe_spinal_indices, ]\nany_severe_spinal_predictions = submission.loc[severe_spinal_indices]\nany_severe_spinal_loss = log_loss(\n    y_true=any_severe_spinal_labels[target_levels],\n    y_pred=any_severe_spinal_predictions[target_levels],\n    sample_weight=any_severe_spinal_weights\n) \n</code></pre>\n<p>Having implemented this part of the metric like this, we would take the first <code>spinal</code> + <code>severe</code> case and its predictions for a specific <code>study_id</code> which in turns serves well for word <code>any</code>.</p>\n<h2>Outro:</h2>\n<hr>\n<p>Thanks for reading it,  I hope it makes sense for you now.<br>\nI find the current metric implementation with <code>.max()</code> - <code>hacky</code> right now. If the real intention to penalize for incorrectly predicted <code>severe</code> label  for a <code>spine canal</code> then the current implementation fails to catch the level component. If the intention to catch any <code>severe</code> label on the entire study for the spinal canal, so why then we have 5 levels to predict (we can mispredict some level and improve on the different one by pure luck)?! </p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I will greatly appreciate your feedback.</p>",
  "messages": [
    {
      "id": "2893483",
      "postDate": "06/27/2024 20:27:28",
      "content": "<p>Hi,</p>\n<h2>Intro</h2>\n<hr>\n<p>I've been experimenting with <code>any_severe_spinal</code> part of the metric for a while and come to the conclusion that there is a chance for <code>random walking</code> of the metric. Let's look at the current metric implementation for that part below:</p>\n<pre><code>any_severe_spinal_labels = pd.Series(\n    solution.loc[solution[] == ].groupby()[].())\nany_severe_spinal_weights = pd.Series(\n    solution.loc[solution[] == ].groupby()[].())\nany_severe_spinal_predictions = pd.Series(\n    submission.loc[submission[] == ].groupby()[].())\nany_severe_spinal_loss = log_loss(\n    y_true=any_severe_spinal_labels,\n    y_pred=any_severe_spinal_predictions,\n    sample_weight=any_severe_spinal_weights\n)\n</code></pre>\n<p>What happens is that by taking <code>.max()</code> we apply it for 5 levels (L1/L2, L2/L3, L3/L4, L4/L5, L5/S1) for a particular <code>study_id</code>. <br>\nWhat if we mispredict some other level as <code>severe</code> with high probability but completely fail to predict the actual level with <code>severe</code> pathology? </p>\n<h2>Real world example</h2>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F4dcfc6ae788d5c151a2637ea909916bb%2FScreenshot%20from%202024-06-27%2015-54-38.png?generation=1719518100343248&amp;alt=media\"></p>\n<p>What we see here is that we failed to predict the <code>severe</code> label for <code>L4/L5</code> level, but <strong>with current implementation it is not a big deal!</strong> - the metric gives us a second chance. How does it happen? Well, the current implementation actually takes <code>.max()</code> which in this case is <code>L3/L4</code> with value <code>~0.52</code>. So now, rather than having <code>true</code> logloss equals to <code>~2.104</code> we get nice <code>~0.6538</code>. </p>\n<h2>Potential ways to fix it:</h2>\n<hr>\n<ol>\n<li>Taking only <code>spinal</code> &amp; <code>severe</code> cases:</li>\n</ol>\n<pre><code>any_severe_spinal_labels = solution.loc[(solution[] == ) &amp; (solution[] == )]\nsevere_spinal_indices = any_severe_spinal_labels.index\nany_severe_spinal_weights = solution.loc[severe_spinal_indices, ]\nany_severe_spinal_predictions = submission.loc[severe_spinal_indices]\nany_severe_spinal_loss = log_loss(\n    y_true=any_severe_spinal_labels[target_levels],\n    y_pred=any_severe_spinal_predictions[target_levels],\n    sample_weight=any_severe_spinal_weights\n) \n</code></pre>\n<p>The problem with that is that we want to take <code>any_severe_spinal</code> observation, where <code>any</code> stands for \"does not matter which level\" - if I understand correctly the host intention.</p>\n<ol>\n<li>Taking the first <code>spinal</code> &amp; <code>severe</code> case per study_id:</li>\n</ol>\n<pre><code>any_severe_spinal_labels = solution.loc[(solution[] == ) &amp; (solution.severe == )]\nany_severe_spinal_labels = any_severe_spinal_labels.drop_duplicates(subset=)\nsevere_spinal_indices = any_severe_spinal_labels.index\nany_severe_spinal_weights = solution.loc[severe_spinal_indices, ]\nany_severe_spinal_predictions = submission.loc[severe_spinal_indices]\nany_severe_spinal_loss = log_loss(\n    y_true=any_severe_spinal_labels[target_levels],\n    y_pred=any_severe_spinal_predictions[target_levels],\n    sample_weight=any_severe_spinal_weights\n) \n</code></pre>\n<p>Having implemented this part of the metric like this, we would take the first <code>spinal</code> + <code>severe</code> case and its predictions for a specific <code>study_id</code> which in turns serves well for word <code>any</code>.</p>\n<h2>Outro:</h2>\n<hr>\n<p>Thanks for reading it,  I hope it makes sense for you now.<br>\nI find the current metric implementation with <code>.max()</code> - <code>hacky</code> right now. If the real intention to penalize for incorrectly predicted <code>severe</code> label  for a <code>spine canal</code> then the current implementation fails to catch the level component. If the intention to catch any <code>severe</code> label on the entire study for the spinal canal, so why then we have 5 levels to predict (we can mispredict some level and improve on the different one by pure luck)?! </p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I will greatly appreciate your feedback.</p>",
      "rawMarkdown": "Hi,\n\n## Intro\n\n---\nI've been experimenting with `any_severe_spinal` part of the metric for a while and come to the conclusion that there is a chance for `random walking` of the metric. Let's look at the current metric implementation for that part below:\n\n\n```python\nany_severe_spinal_labels = pd.Series(\n    solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['severe'].max())\nany_severe_spinal_weights = pd.Series(\n    solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['sample_weight'].max())\nany_severe_spinal_predictions = pd.Series(\n    submission.loc[submission['condition'] == 'spinal'].groupby('study_id')['severe'].max())\nany_severe_spinal_loss = log_loss(\n    y_true=any_severe_spinal_labels,\n    y_pred=any_severe_spinal_predictions,\n    sample_weight=any_severe_spinal_weights\n)\n```\n\nWhat happens is that by taking `.max()` we apply it for 5 levels (L1/L2, L2/L3, L3/L4, L4/L5, L5/S1) for a particular `study_id`. \nWhat if we mispredict some other level as `severe` with high probability but completely fail to predict the actual level with `severe` pathology? \n\n##Real world example\n\n---\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F4dcfc6ae788d5c151a2637ea909916bb%2FScreenshot%20from%202024-06-27%2015-54-38.png?generation=1719518100343248&alt=media)\n\nWhat we see here is that we failed to predict the `severe` label for `L4/L5` level, but **with current implementation it is not a big deal!** - the metric gives us a second chance. How does it happen? Well, the current implementation actually takes `.max()` which in this case is `L3/L4` with value `~0.52`. So now, rather than having `true` logloss equals to `~2.104` we get nice `~0.6538`. \n\n##Potential ways to fix it:\n\n---\n\n1. Taking only `spinal` & `severe` cases:\n```python\nany_severe_spinal_labels = solution.loc[(solution['condition'] == 'spinal') & (solution['severe'] == 1)]\nsevere_spinal_indices = any_severe_spinal_labels.index\nany_severe_spinal_weights = solution.loc[severe_spinal_indices, 'sample_weight']\nany_severe_spinal_predictions = submission.loc[severe_spinal_indices]\nany_severe_spinal_loss = log_loss(\n    y_true=any_severe_spinal_labels[target_levels],\n    y_pred=any_severe_spinal_predictions[target_levels],\n    sample_weight=any_severe_spinal_weights\n) \n```\nThe problem with that is that we want to take `any_severe_spinal` observation, where `any` stands for \"does not matter which level\" - if I understand correctly the host intention.\n\n2. Taking the first `spinal` & `severe` case per study_id:\n```python\nany_severe_spinal_labels = solution.loc[(solution['condition'] == 'spinal') & (solution.severe == 1)]\nany_severe_spinal_labels = any_severe_spinal_labels.drop_duplicates(subset='study_id')\nsevere_spinal_indices = any_severe_spinal_labels.index\nany_severe_spinal_weights = solution.loc[severe_spinal_indices, 'sample_weight']\nany_severe_spinal_predictions = submission.loc[severe_spinal_indices]\nany_severe_spinal_loss = log_loss(\n    y_true=any_severe_spinal_labels[target_levels],\n    y_pred=any_severe_spinal_predictions[target_levels],\n    sample_weight=any_severe_spinal_weights\n) \n```\nHaving implemented this part of the metric like this, we would take the first `spinal` + `severe` case and its predictions for a specific `study_id` which in turns serves well for word `any`.\n\n##Outro:\n\n---\nThanks for reading it,  I hope it makes sense for you now.\nI find the current metric implementation with `.max()` - `hacky` right now. If the real intention to penalize for incorrectly predicted `severe` label  for a `spine canal` then the current implementation fails to catch the level component. If the intention to catch any `severe` label on the entire study for the spinal canal, so why then we have 5 levels to predict (we can mispredict some level and improve on the different one by pure luck)?! \n\n@sohier I will greatly appreciate your feedback.",
      "votes": null
    },
    {
      "id": "2893595",
      "postDate": "06/27/2024 22:23:29",
      "content": "<p>I also raised this <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/510363#2878218\" target=\"_blank\">here</a>. I think it is likely that the behaviour of the metric is intended given the  <code>_any</code> prefix in the calculation.</p>",
      "rawMarkdown": "I also raised this [here](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/510363#2878218). I think it is likely that the behaviour of the metric is intended given the  `_any` prefix in the calculation.",
      "votes": null
    },
    {
      "id": "2893632",
      "postDate": "06/27/2024 23:49:12",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> yes, thanks, actually I checked that before posting the topic. I was not sure about the order thing (agreed on that) and if we were on the same page. Idk if <code>any</code> prefix makes the metric a voodo thing or the second logloss implementation itself.</p>",
      "rawMarkdown": "brendanartley yes, thanks, actually I checked that before posting the topic. I was not sure about the order thing (agreed on that) and if we were on the same page. Idk if `any` prefix makes the metric a voodo thing or the second logloss implementation itself.",
      "votes": null
    },
    {
      "id": "2895072",
      "postDate": "06/28/2024 22:44:01",
      "content": "<p>I don't think I follow your concern. From the patient's perspective the overall diagnosis matters more than the localization, which can be corrected as long as they remain in contact with the medical system.</p>",
      "rawMarkdown": "I don't think I follow your concern. From the patient's perspective the overall diagnosis matters more than the localization, which can be corrected as long as they remain in contact with the medical system.",
      "votes": null
    },
    {
      "id": "2897781",
      "postDate": "06/30/2024 17:40:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>Why \"any_severe\" calculations are based on just  \"spinal\" condition which are 5-labels, wouldn't it be in best interest of patients if we consider 'any_severe\" calculations based on all 25-labels (\"spinal_canal_stenosis\" :5, \"neural_foraminal_narrowing\":10,  \"subarticular_stenosis\":10) ?</p>\n<p>Are conditions \"neural_foraminal_narrowing\"  \"subarticular_stenosis\" not as important as \"spinal_canal_stenosis\" ?</p>",
      "rawMarkdown": "Hi @sohier,\n\nWhy \"any_severe\" calculations are based on just  \"spinal\" condition which are 5-labels, wouldn't it be in best interest of patients if we consider 'any_severe\" calculations based on all 25-labels (\"spinal_canal_stenosis\" :5, \"neural_foraminal_narrowing\":10,  \"subarticular_stenosis\":10) ?\n\nAre conditions \"neural_foraminal_narrowing\"  \"subarticular_stenosis\" not as important as \"spinal_canal_stenosis\" ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2893595,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "06/27/2024 22:23:29",
      "content": "<p>I also raised this <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/510363#2878218\" target=\"_blank\">here</a>. I think it is likely that the behaviour of the metric is intended given the  <code>_any</code> prefix in the calculation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2893632,
          "author_name": "sergiosaharovskiy",
          "author_url": "",
          "post_date": "06/27/2024 23:49:12",
          "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> yes, thanks, actually I checked that before posting the topic. I was not sure about the order thing (agreed on that) and if we were on the same page. Idk if <code>any</code> prefix makes the metric a voodo thing or the second logloss implementation itself.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2895072,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "06/28/2024 22:44:01",
      "content": "<p>I don't think I follow your concern. From the patient's perspective the overall diagnosis matters more than the localization, which can be corrected as long as they remain in contact with the medical system.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2897781,
          "author_name": "rohitchaudhari25",
          "author_url": "",
          "post_date": "06/30/2024 17:40:17",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>Why \"any_severe\" calculations are based on just  \"spinal\" condition which are 5-labels, wouldn't it be in best interest of patients if we consider 'any_severe\" calculations based on all 25-labels (\"spinal_canal_stenosis\" :5, \"neural_foraminal_narrowing\":10,  \"subarticular_stenosis\":10) ?</p>\n<p>Are conditions \"neural_foraminal_narrowing\"  \"subarticular_stenosis\" not as important as \"spinal_canal_stenosis\" ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2893483": "Hi,\n\n## Intro\n\n---\nI've been experimenting with `any_severe_spinal` part of the metric for a while and come to the conclusion that there is a chance for `random walking` of the metric. Let's look at the current metric implementation for that part below:\n\n\n```python\nany_severe_spinal_labels = pd.Series(\n    solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['severe'].max())\nany_severe_spinal_weights = pd.Series(\n    solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['sample_weight'].max())\nany_severe_spinal_predictions = pd.Series(\n    submission.loc[submission['condition'] == 'spinal'].groupby('study_id')['severe'].max())\nany_severe_spinal_loss = log_loss(\n    y_true=any_severe_spinal_labels,\n    y_pred=any_severe_spinal_predictions,\n    sample_weight=any_severe_spinal_weights\n)\n```\n\nWhat happens is that by taking `.max()` we apply it for 5 levels (L1/L2, L2/L3, L3/L4, L4/L5, L5/S1) for a particular `study_id`. \nWhat if we mispredict some other level as `severe` with high probability but completely fail to predict the actual level with `severe` pathology? \n\n##Real world example\n\n---\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F4dcfc6ae788d5c151a2637ea909916bb%2FScreenshot%20from%202024-06-27%2015-54-38.png?generation=1719518100343248&alt=media)\n\nWhat we see here is that we failed to predict the `severe` label for `L4/L5` level, but **with current implementation it is not a big deal!** - the metric gives us a second chance. How does it happen? Well, the current implementation actually takes `.max()` which in this case is `L3/L4` with value `~0.52`. So now, rather than having `true` logloss equals to `~2.104` we get nice `~0.6538`. \n\n##Potential ways to fix it:\n\n---\n\n1. Taking only `spinal` & `severe` cases:\n```python\nany_severe_spinal_labels = solution.loc[(solution['condition'] == 'spinal') & (solution['severe'] == 1)]\nsevere_spinal_indices = any_severe_spinal_labels.index\nany_severe_spinal_weights = solution.loc[severe_spinal_indices, 'sample_weight']\nany_severe_spinal_predictions = submission.loc[severe_spinal_indices]\nany_severe_spinal_loss = log_loss(\n    y_true=any_severe_spinal_labels[target_levels],\n    y_pred=any_severe_spinal_predictions[target_levels],\n    sample_weight=any_severe_spinal_weights\n) \n```\nThe problem with that is that we want to take `any_severe_spinal` observation, where `any` stands for \"does not matter which level\" - if I understand correctly the host intention.\n\n2. Taking the first `spinal` & `severe` case per study_id:\n```python\nany_severe_spinal_labels = solution.loc[(solution['condition'] == 'spinal') & (solution.severe == 1)]\nany_severe_spinal_labels = any_severe_spinal_labels.drop_duplicates(subset='study_id')\nsevere_spinal_indices = any_severe_spinal_labels.index\nany_severe_spinal_weights = solution.loc[severe_spinal_indices, 'sample_weight']\nany_severe_spinal_predictions = submission.loc[severe_spinal_indices]\nany_severe_spinal_loss = log_loss(\n    y_true=any_severe_spinal_labels[target_levels],\n    y_pred=any_severe_spinal_predictions[target_levels],\n    sample_weight=any_severe_spinal_weights\n) \n```\nHaving implemented this part of the metric like this, we would take the first `spinal` + `severe` case and its predictions for a specific `study_id` which in turns serves well for word `any`.\n\n##Outro:\n\n---\nThanks for reading it,  I hope it makes sense for you now.\nI find the current metric implementation with `.max()` - `hacky` right now. If the real intention to penalize for incorrectly predicted `severe` label  for a `spine canal` then the current implementation fails to catch the level component. If the intention to catch any `severe` label on the entire study for the spinal canal, so why then we have 5 levels to predict (we can mispredict some level and improve on the different one by pure luck)?! \n\n@sohier I will greatly appreciate your feedback.",
    "2893595": "I also raised this [here](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/510363#2878218). I think it is likely that the behaviour of the metric is intended given the  `_any` prefix in the calculation.",
    "2893632": "brendanartley yes, thanks, actually I checked that before posting the topic. I was not sure about the order thing (agreed on that) and if we were on the same page. Idk if `any` prefix makes the metric a voodo thing or the second logloss implementation itself.",
    "2895072": "I don't think I follow your concern. From the patient's perspective the overall diagnosis matters more than the localization, which can be corrected as long as they remain in contact with the medical system.",
    "2897781": "Hi @sohier,\n\nWhy \"any_severe\" calculations are based on just  \"spinal\" condition which are 5-labels, wouldn't it be in best interest of patients if we consider 'any_severe\" calculations based on all 25-labels (\"spinal_canal_stenosis\" :5, \"neural_foraminal_narrowing\":10,  \"subarticular_stenosis\":10) ?\n\nAre conditions \"neural_foraminal_narrowing\"  \"subarticular_stenosis\" not as important as \"spinal_canal_stenosis\" ?"
  },
  "source": "meta"
}