{
  "id": 511055,
  "title": "Competition Metric Behavior vs Incorrect predictions",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/511055",
  "author_name": "SSS",
  "post_date": "2024-06-08T21:11:58.142000",
  "votes": 30,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I have experimented with the metric a bit trying to model its behavior.<br>\nOn the high level we can look at the <a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549\" target=\"_blank\"><strong>metric implementation</strong></a> and get a grasp of what is going on.  Additionally to the code, I'd like to share few charts:</p>\n<h3>1. Metric penalty</h3>\n<p>I have sampled 30 examples of each label and incrementally replaced the true label with the wrong one for each condition separately keeping everything else untouched as follows: </p>\n<pre><code>normal_sample   = solution.query().head()\nsevere_sample   = solution.query().head()\nmoderate_sample = solution.query().head()\n\nsolution_sample   = pd.concat([normal_sample, severe_sample, moderate_sample], ignore_index=).copy()\nsubmission_sample = solution_sample.drop(columns=).copy()\n</code></pre>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F0e9dfbfa554aba238c1fecd9c06b0432%2Fscore_understanding.png?generation=1717879406469223&amp;alt=media\"></p>\n<p><strong>Notes</strong>:<br>\nHow to read it: out of 30 <code>severe</code> cases in the 90 total observations if we incorrectly predict all 30 we get the logloss around 24 (keeping other cases untouched).</p>\n<ol>\n<li>Obviously metric punishes the most for the incorrectly predicted <code>severe</code> targets. The metric punishes even more when we fail to predict  <code>severe</code>  for <code>spinal_canal</code>.</li>\n<li>The penalty for the incorrectly predicted labels is linear (thanks captain), though, not exactly proportional to the competition weights [1, 2, 4].</li>\n</ol>\n<h3>2. Train hill climbing (coordinates descent)</h3>\n<p>We can play with hill climbing based on the train set and see how the score changes.</p>\n<p>Example (the best score I have got by having <code>[0.55, 0.20, 0.25]</code>):</p>\n<table>\n<thead>\n<tr>\n<th>normal_mild</th>\n<th>moderate</th>\n<th>severe</th>\n<th>score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.55</td>\n<td>0.20</td>\n<td>0.25</td>\n<td>0.910161</td>\n</tr>\n<tr>\n<td>0.60</td>\n<td>0.20</td>\n<td>0.20</td>\n<td>0.915495</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p>Let's visualize it:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F762e554fa6be8af9c017bdb250d84a03%2FScreencastfrom06-08-2024035639PM-ezgif.com-video-to-gif-converter.gif?generation=1717880655774457&amp;alt=media\"></p>\n<p>Hope it makes more sense for you now. It actually correlates with the leaderboard probing:</p>\n<table>\n<thead>\n<tr>\n<th>Weights</th>\n<th>Score Before Patch v2</th>\n<th>Score After Patch v2 (Current)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Frequencies</td>\n<td>1.24</td>\n<td>1.21</td>\n</tr>\n<tr>\n<td>[0.33, 0.33, 0.33]</td>\n<td>0.89</td>\n<td>1.02</td>\n</tr>\n<tr>\n<td>[0.424223, 0.303108, 0.272669]</td>\n<td>0.84</td>\n<td>0.98</td>\n</tr>\n<tr>\n<td>[0.5, 0.15, 0.35]  (hill climbing)</td>\n<td>0.97</td>\n<td>1.25</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 2862656,
      "postDate": "2024-06-08T21:11:58.143Z",
      "content": "<p>Hi,</p>\n<p>I have experimented with the metric a bit trying to model its behavior.<br>\nOn the high level we can look at the <a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549\" target=\"_blank\"><strong>metric implementation</strong></a> and get a grasp of what is going on.  Additionally to the code, I'd like to share few charts:</p>\n<h3>1. Metric penalty</h3>\n<p>I have sampled 30 examples of each label and incrementally replaced the true label with the wrong one for each condition separately keeping everything else untouched as follows: </p>\n<pre><code>normal_sample   = solution.query().head()\nsevere_sample   = solution.query().head()\nmoderate_sample = solution.query().head()\n\nsolution_sample   = pd.concat([normal_sample, severe_sample, moderate_sample], ignore_index=).copy()\nsubmission_sample = solution_sample.drop(columns=).copy()\n</code></pre>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F0e9dfbfa554aba238c1fecd9c06b0432%2Fscore_understanding.png?generation=1717879406469223&amp;alt=media\"></p>\n<p><strong>Notes</strong>:<br>\nHow to read it: out of 30 <code>severe</code> cases in the 90 total observations if we incorrectly predict all 30 we get the logloss around 24 (keeping other cases untouched).</p>\n<ol>\n<li>Obviously metric punishes the most for the incorrectly predicted <code>severe</code> targets. The metric punishes even more when we fail to predict  <code>severe</code>  for <code>spinal_canal</code>.</li>\n<li>The penalty for the incorrectly predicted labels is linear (thanks captain), though, not exactly proportional to the competition weights [1, 2, 4].</li>\n</ol>\n<h3>2. Train hill climbing (coordinates descent)</h3>\n<p>We can play with hill climbing based on the train set and see how the score changes.</p>\n<p>Example (the best score I have got by having <code>[0.55, 0.20, 0.25]</code>):</p>\n<table>\n<thead>\n<tr>\n<th>normal_mild</th>\n<th>moderate</th>\n<th>severe</th>\n<th>score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.55</td>\n<td>0.20</td>\n<td>0.25</td>\n<td>0.910161</td>\n</tr>\n<tr>\n<td>0.60</td>\n<td>0.20</td>\n<td>0.20</td>\n<td>0.915495</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p>Let's visualize it:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F762e554fa6be8af9c017bdb250d84a03%2FScreencastfrom06-08-2024035639PM-ezgif.com-video-to-gif-converter.gif?generation=1717880655774457&amp;alt=media\"></p>\n<p>Hope it makes more sense for you now. It actually correlates with the leaderboard probing:</p>\n<table>\n<thead>\n<tr>\n<th>Weights</th>\n<th>Score Before Patch v2</th>\n<th>Score After Patch v2 (Current)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Frequencies</td>\n<td>1.24</td>\n<td>1.21</td>\n</tr>\n<tr>\n<td>[0.33, 0.33, 0.33]</td>\n<td>0.89</td>\n<td>1.02</td>\n</tr>\n<tr>\n<td>[0.424223, 0.303108, 0.272669]</td>\n<td>0.84</td>\n<td>0.98</td>\n</tr>\n<tr>\n<td>[0.5, 0.15, 0.35]  (hill climbing)</td>\n<td>0.97</td>\n<td>1.25</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Hi,\n\nI have experimented with the metric a bit trying to model its behavior.\nOn the high level we can look at the [**metric implementation**](https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549) and get a grasp of what is going on.  Additionally to the code, I'd like to share few charts:\n\n### 1. Metric penalty\nI have sampled 30 examples of each label and incrementally replaced the true label with the wrong one for each condition separately keeping everything else untouched as follows: \n\n```python\nnormal_sample   = solution.query('normal_mild == 1').head(30)\nsevere_sample   = solution.query('severe == 1').head(30)\nmoderate_sample = solution.query('moderate == 1').head(30)\n\nsolution_sample   = pd.concat([normal_sample, severe_sample, moderate_sample], ignore_index=True).copy()\nsubmission_sample = solution_sample.drop(columns='sample_weight').copy()\n```\n---\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F0e9dfbfa554aba238c1fecd9c06b0432%2Fscore_understanding.png?generation=1717879406469223&alt=media)\n\n**Notes**:\nHow to read it: out of 30 `severe` cases in the 90 total observations if we incorrectly predict all 30 we get the logloss around 24 (keeping other cases untouched).\n\n1.  Obviously metric punishes the most for the incorrectly predicted `severe` targets. The metric punishes even more when we fail to predict  `severe`  for `spinal_canal`.\n2. The penalty for the incorrectly predicted labels is linear (thanks captain), though, not exactly proportional to the competition weights [1, 2, 4].\n\n### 2. Train hill climbing (coordinates descent)\nWe can play with hill climbing based on the train set and see how the score changes.\n\n\nExample (the best score I have got by having `[0.55, 0.20, 0.25]`):\n| normal_mild | moderate | severe |   score  |\n|-------------|----------|--------|----------|\n|     0.55    |   0.20   |  0.25  | 0.910161 |\n|     0.60    |   0.20   |  0.20  | 0.915495 |\n\n\n---\nLet's visualize it:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F762e554fa6be8af9c017bdb250d84a03%2FScreencastfrom06-08-2024035639PM-ezgif.com-video-to-gif-converter.gif?generation=1717880655774457&alt=media)\n\nHope it makes more sense for you now. It actually correlates with the leaderboard probing:\n| Weights                      | Score Before Patch v2 | Score After Patch v2 (Current) |\n|------------------------------|----------------------|---------------------|\n| Frequencies                  | 1.24                 | 1.21                |\n| [0.33, 0.33, 0.33]           | 0.89                 | 1.02                |\n| [0.424223, 0.303108, 0.272669] | 0.84                 | 0.98                |\n| [0.5, 0.15, 0.35]  (hill climbing)          | 0.97                 | 1.25                |",
      "votes": 30
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2862656": "Hi,\n\nI have experimented with the metric a bit trying to model its behavior.\nOn the high level we can look at the [**metric implementation**](https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549) and get a grasp of what is going on.  Additionally to the code, I'd like to share few charts:\n\n### 1. Metric penalty\nI have sampled 30 examples of each label and incrementally replaced the true label with the wrong one for each condition separately keeping everything else untouched as follows: \n\n```python\nnormal_sample   = solution.query('normal_mild == 1').head(30)\nsevere_sample   = solution.query('severe == 1').head(30)\nmoderate_sample = solution.query('moderate == 1').head(30)\n\nsolution_sample   = pd.concat([normal_sample, severe_sample, moderate_sample], ignore_index=True).copy()\nsubmission_sample = solution_sample.drop(columns='sample_weight').copy()\n```\n---\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F0e9dfbfa554aba238c1fecd9c06b0432%2Fscore_understanding.png?generation=1717879406469223&alt=media)\n\n**Notes**:\nHow to read it: out of 30 `severe` cases in the 90 total observations if we incorrectly predict all 30 we get the logloss around 24 (keeping other cases untouched).\n\n1.  Obviously metric punishes the most for the incorrectly predicted `severe` targets. The metric punishes even more when we fail to predict  `severe`  for `spinal_canal`.\n2. The penalty for the incorrectly predicted labels is linear (thanks captain), though, not exactly proportional to the competition weights [1, 2, 4].\n\n### 2. Train hill climbing (coordinates descent)\nWe can play with hill climbing based on the train set and see how the score changes.\n\n\nExample (the best score I have got by having `[0.55, 0.20, 0.25]`):\n| normal_mild | moderate | severe |   score  |\n|-------------|----------|--------|----------|\n|     0.55    |   0.20   |  0.25  | 0.910161 |\n|     0.60    |   0.20   |  0.20  | 0.915495 |\n\n\n---\nLet's visualize it:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F762e554fa6be8af9c017bdb250d84a03%2FScreencastfrom06-08-2024035639PM-ezgif.com-video-to-gif-converter.gif?generation=1717880655774457&alt=media)\n\nHope it makes more sense for you now. It actually correlates with the leaderboard probing:\n| Weights                      | Score Before Patch v2 | Score After Patch v2 (Current) |\n|------------------------------|----------------------|---------------------|\n| Frequencies                  | 1.24                 | 1.21                |\n| [0.33, 0.33, 0.33]           | 0.89                 | 1.02                |\n| [0.424223, 0.303108, 0.272669] | 0.84                 | 0.98                |\n| [0.5, 0.15, 0.35]  (hill climbing)          | 0.97                 | 1.25                |"
  }
}