{
  "id": 510363,
  "title": "Second metric patch and rescore",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/510363",
  "author_name": "Sohier Dane",
  "post_date": "2024-06-05T22:21:00.405000",
  "votes": 17,
  "comment_count": 6,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/509034#2849168\" target=\"_blank\">identified an issue with the metric</a> that I missed in <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522\" target=\"_blank\">my initial patch</a>. Essentially, I failed to fully propagate all of the necessary changes when migrating from a pre-launch version of the dataset where the submission and solution file used a wide rather than long format.</p>\n<p>The update to the metric is quite simple and is already live. I will be rescoring all existing submissions shortly. That process should take less than an hour.</p>\n<p>The metric shouldn't have required multiple patches. I know the bugs have been disruptive and apologize for allowing them to get deployed.</p>\n<p>Edit: The rescore is now complete.</p>",
  "messages": [
    {
      "id": 2857468,
      "postDate": "2024-06-05T22:21:00.407Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/509034#2849168\" target=\"_blank\">identified an issue with the metric</a> that I missed in <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522\" target=\"_blank\">my initial patch</a>. Essentially, I failed to fully propagate all of the necessary changes when migrating from a pre-launch version of the dataset where the submission and solution file used a wide rather than long format.</p>\n<p>The update to the metric is quite simple and is already live. I will be rescoring all existing submissions shortly. That process should take less than an hour.</p>\n<p>The metric shouldn't have required multiple patches. I know the bugs have been disruptive and apologize for allowing them to get deployed.</p>\n<p>Edit: The rescore is now complete.</p>",
      "rawMarkdown": "@vaillant [identified an issue with the metric](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/509034#2849168) that I missed in [my initial patch](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522). Essentially, I failed to fully propagate all of the necessary changes when migrating from a pre-launch version of the dataset where the submission and solution file used a wide rather than long format.\n\nThe update to the metric is quite simple and is already live. I will be rescoring all existing submissions shortly. That process should take less than an hour.\n\nThe metric shouldn't have required multiple patches. I know the bugs have been disruptive and apologize for allowing them to get deployed.\n\nEdit: The rescore is now complete.",
      "votes": 17
    },
    {
      "id": 2858539,
      "postDate": "2024-06-06T14:40:24.380Z",
      "content": "<p>I will leave it here.</p>\n<pre><code>\ncondition_weights.append( / solution.loc[condition_indices, ].nunique()) &gt;&gt; condition_weights.append()\n</code></pre>\n<table>\n<thead>\n<tr>\n<th>Weights</th>\n<th>Score Before Patch v2</th>\n<th>Score After Patch v2 (Current)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Frequencies</td>\n<td>1.24</td>\n<td>1.21</td>\n</tr>\n<tr>\n<td>[0.33, 0.33, 0.33]</td>\n<td>0.89</td>\n<td>1.02</td>\n</tr>\n<tr>\n<td>[0.424223, 0.303108, 0.272669]</td>\n<td>0.84</td>\n<td>0.98</td>\n</tr>\n<tr>\n<td>[0.5, 0.15, 0.35]  (hill climbing)</td>\n<td>0.97</td>\n<td>1.25</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "I will leave it here.\n\n```python\n# Change patch v2\ncondition_weights.append(1 / solution.loc[condition_indices, 'location'].nunique()) >> condition_weights.append(1)\n```\n\n| Weights                      | Score Before Patch v2 | Score After Patch v2 (Current) |\n|------------------------------|----------------------|---------------------|\n| Frequencies                  | 1.24                 | 1.21                |\n| [0.33, 0.33, 0.33]           | 0.89                 | 1.02                |\n| [0.424223, 0.303108, 0.272669] | 0.84                 | 0.98                |\n| [0.5, 0.15, 0.35]  (hill climbing)          | 0.97                 | 1.25                |",
      "votes": 3,
      "replies": [
        {
          "id": 2864550,
          "postDate": "2024-06-10T07:49:32.133Z",
          "content": "<p>I am very curious.</p>\n<p>1) Your score is 1.21. So how do you know that [0.33, 0.33, 0.33] will produce     1.02 and [0.424223, 0.303108, 0.272669] will produce 0.98? One way to do this is to submit. But your score would have been much higher if you had done that.</p>\n<p>2) When I tried these values the score I got was different it was '1' for [0.33, 0.33, 0.33] and '2' for [0.424223, 0.303108, 0.272669].</p>\n<p>Your response is highly encouraged and will be very valuable. Thankyou.</p>",
          "rawMarkdown": "I am very curious.\n\n1) Your score is 1.21. So how do you know that [0.33, 0.33, 0.33] will produce \t1.02 and [0.424223, 0.303108, 0.272669] will produce 0.98? One way to do this is to submit. But your score would have been much higher if you had done that.\n\n2) When I tried these values the score I got was different it was '1' for [0.33, 0.33, 0.33] and '2' for [0.424223, 0.303108, 0.272669].\n\nYour response is highly encouraged and will be very valuable. Thankyou.",
          "replies": [
            {
              "id": 2864723,
              "postDate": "2024-06-10T09:44:16.623Z",
              "content": "<p>This difference is due to rescoring because of metric change. I also previously submitted a baseline with [0.424223, 0.303108, 0.272669] and got 0.98. But now the same submission score has been updated 1.0.</p>",
              "rawMarkdown": "This difference is due to rescoring because of metric change. I also previously submitted a baseline with [0.424223, 0.303108, 0.272669] and got 0.98. But now the same submission score has been updated 1.0."
            }
          ]
        }
      ]
    },
    {
      "id": 2878218,
      "postDate": "2024-06-18T19:46:10.603Z",
      "content": "<p>The order of the predictions does not seem to matter in the second log loss calculation. This may be intended given the \"any\" prefix, but thought I would share just in case. <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>For example, if you add this code block just before the line that starts with <code>any_severe_spinal_labels = ..</code> in <a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549\" target=\"_blank\">the competition metric notebook</a>, the final score is always the same.</p>\n<pre><code>    \n    = submission[].values.tolist()\n    =\n     i in range(, len(vals), ):\n        = random.sample(vals[i:i+], k=)\n        .extend(new_order)\n    []= arr\n</code></pre>",
      "rawMarkdown": "The order of the predictions does not seem to matter in the second log loss calculation. This may be intended given the \"any\" prefix, but thought I would share just in case. @sohier \n\nFor example, if you add this code block just before the line that starts with `any_severe_spinal_labels = ..` in [the competition metric notebook](https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549), the final score is always the same.\n\n```\n    # Randomly shuffle every 5 values\n    vals= submission[\"severe\"].values.tolist()\n    arr= []\n    for i in range(0, len(vals), 5):\n        new_order= random.sample(vals[i:i+5], k=5)\n        arr.extend(new_order)\n    submission[\"severe\"]= arr\n```"
    },
    {
      "id": 2862512,
      "postDate": "2024-06-08T19:29:37.067Z",
      "content": "<p>How can we accurately identify and annotate the correct MRI image slices from a test set of lumbar spine MRI scans, in order to predict the severity of various conditions at specific vertebral levels (L1-S1), given that the test set does not provide explicit pointers</p>",
      "rawMarkdown": "How can we accurately identify and annotate the correct MRI image slices from a test set of lumbar spine MRI scans, in order to predict the severity of various conditions at specific vertebral levels (L1-S1), given that the test set does not provide explicit pointers"
    },
    {
      "id": 2894027,
      "postDate": "2024-06-28T07:39:26.940Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2858539,
      "author_name": "SSS",
      "author_url": "",
      "post_date": "2024-06-06T14:40:24.380000",
      "content": "<p>I will leave it here.</p>\n<pre><code>\ncondition_weights.append( / solution.loc[condition_indices, ].nunique()) &gt;&gt; condition_weights.append()\n</code></pre>\n<table>\n<thead>\n<tr>\n<th>Weights</th>\n<th>Score Before Patch v2</th>\n<th>Score After Patch v2 (Current)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Frequencies</td>\n<td>1.24</td>\n<td>1.21</td>\n</tr>\n<tr>\n<td>[0.33, 0.33, 0.33]</td>\n<td>0.89</td>\n<td>1.02</td>\n</tr>\n<tr>\n<td>[0.424223, 0.303108, 0.272669]</td>\n<td>0.84</td>\n<td>0.98</td>\n</tr>\n<tr>\n<td>[0.5, 0.15, 0.35]  (hill climbing)</td>\n<td>0.97</td>\n<td>1.25</td>\n</tr>\n</tbody>\n</table>",
      "votes": 3,
      "replies": [
        {
          "id": 2864550,
          "author_name": "Devsya ",
          "author_url": "",
          "post_date": "2024-06-10T07:49:32.133000",
          "content": "<p>I am very curious.</p>\n<p>1) Your score is 1.21. So how do you know that [0.33, 0.33, 0.33] will produce     1.02 and [0.424223, 0.303108, 0.272669] will produce 0.98? One way to do this is to submit. But your score would have been much higher if you had done that.</p>\n<p>2) When I tried these values the score I got was different it was '1' for [0.33, 0.33, 0.33] and '2' for [0.424223, 0.303108, 0.272669].</p>\n<p>Your response is highly encouraged and will be very valuable. Thankyou.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2864723,
              "author_name": "coderRKJ",
              "author_url": "",
              "post_date": "2024-06-10T09:44:16.623000",
              "content": "<p>This difference is due to rescoring because of metric change. I also previously submitted a baseline with [0.424223, 0.303108, 0.272669] and got 0.98. But now the same submission score has been updated 1.0.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2878218,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2024-06-18T19:46:10.603000",
      "content": "<p>The order of the predictions does not seem to matter in the second log loss calculation. This may be intended given the \"any\" prefix, but thought I would share just in case. <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>For example, if you add this code block just before the line that starts with <code>any_severe_spinal_labels = ..</code> in <a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549\" target=\"_blank\">the competition metric notebook</a>, the final score is always the same.</p>\n<pre><code>    \n    = submission[].values.tolist()\n    =\n     i in range(, len(vals), ):\n        = random.sample(vals[i:i+], k=)\n        .extend(new_order)\n    []= arr\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2862512,
      "author_name": "Exalted Joseph",
      "author_url": "",
      "post_date": "2024-06-08T19:29:37.067000",
      "content": "<p>How can we accurately identify and annotate the correct MRI image slices from a test set of lumbar spine MRI scans, in order to predict the severity of various conditions at specific vertebral levels (L1-S1), given that the test set does not provide explicit pointers</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2894027,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-28T07:39:26.940000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2857468": "@vaillant [identified an issue with the metric](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/509034#2849168) that I missed in [my initial patch](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522). Essentially, I failed to fully propagate all of the necessary changes when migrating from a pre-launch version of the dataset where the submission and solution file used a wide rather than long format.\n\nThe update to the metric is quite simple and is already live. I will be rescoring all existing submissions shortly. That process should take less than an hour.\n\nThe metric shouldn't have required multiple patches. I know the bugs have been disruptive and apologize for allowing them to get deployed.\n\nEdit: The rescore is now complete.",
    "2858539": "I will leave it here.\n\n```python\n# Change patch v2\ncondition_weights.append(1 / solution.loc[condition_indices, 'location'].nunique()) >> condition_weights.append(1)\n```\n\n| Weights                      | Score Before Patch v2 | Score After Patch v2 (Current) |\n|------------------------------|----------------------|---------------------|\n| Frequencies                  | 1.24                 | 1.21                |\n| [0.33, 0.33, 0.33]           | 0.89                 | 1.02                |\n| [0.424223, 0.303108, 0.272669] | 0.84                 | 0.98                |\n| [0.5, 0.15, 0.35]  (hill climbing)          | 0.97                 | 1.25                |",
    "2878218": "The order of the predictions does not seem to matter in the second log loss calculation. This may be intended given the \"any\" prefix, but thought I would share just in case. @sohier \n\nFor example, if you add this code block just before the line that starts with `any_severe_spinal_labels = ..` in [the competition metric notebook](https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549), the final score is always the same.\n\n```\n    # Randomly shuffle every 5 values\n    vals= submission[\"severe\"].values.tolist()\n    arr= []\n    for i in range(0, len(vals), 5):\n        new_order= random.sample(vals[i:i+5], k=5)\n        arr.extend(new_order)\n    submission[\"severe\"]= arr\n```",
    "2862512": "How can we accurately identify and annotate the correct MRI image slices from a test set of lumbar spine MRI scans, in order to predict the severity of various conditions at specific vertebral levels (L1-S1), given that the test set does not provide explicit pointers",
    "2894027": ""
  }
}