{
  "id": 513204,
  "title": "What is exactly the sample_weight in the metric score?",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/513204",
  "author_name": "",
  "post_date": "2024-06-19T03:46:15.852136500Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I made successful submissions, I pass this to latest scoring funciton, and it's looking for a mysterious column</p>\n<p>sample_weight=solution.loc[condition_indices, 'sample_weight'].values</p>\n<p>neither the Submission or the Solution supposed to have a sample_weight?</p>\n<p>How is this calculated?</p>\n<p>Submission<br>\nrow_id    normal_mild moderate    severe<br>\n0    0000_left_neural_foraminal_narrowing_l1_l2  0.958009    0.040073    0.000177<br>\n1    0000_left_neural_foraminal_narrowing_l2_l3  0.889687    0.101157    0.003351<br>\n2    0000_left_neural_foraminal_narrowing_l3_l4  0.701213    0.266386    0.023497<br>\n3    0000_left_neural_foraminal_narrowing_l4_l5  0.510915    0.391577    0.092108<br>\n4    0000_left_neural_foraminal_narrowing_l5_s1  0.557290    0.308225    0.127156</p>\n<p>Solution<br>\nrow_id    normal_mild moderate    severe<br>\n0    0000_left_neural_foraminal_narrowing_l1_l2  1.0 0.0 0.0<br>\n1    0000_left_neural_foraminal_narrowing_l2_l3  0.0 1.0 0.0<br>\n2    0000_left_neural_foraminal_narrowing_l3_l4  0.0 1.0 0.0<br>\n3    0000_left_neural_foraminal_narrowing_l4_l5  0.0 0.0 1.0<br>\n4    0000_left_neural_foraminal_narrowing_l5_s1  0.0 1.0 0.0</p>",
  "messages": [
    {
      "id": "2878545",
      "postDate": "06/19/2024 03:46:15",
      "content": "<p>I made successful submissions, I pass this to latest scoring funciton, and it's looking for a mysterious column</p>\n<p>sample_weight=solution.loc[condition_indices, 'sample_weight'].values</p>\n<p>neither the Submission or the Solution supposed to have a sample_weight?</p>\n<p>How is this calculated?</p>\n<p>Submission<br>\nrow_id    normal_mild moderate    severe<br>\n0    0000_left_neural_foraminal_narrowing_l1_l2  0.958009    0.040073    0.000177<br>\n1    0000_left_neural_foraminal_narrowing_l2_l3  0.889687    0.101157    0.003351<br>\n2    0000_left_neural_foraminal_narrowing_l3_l4  0.701213    0.266386    0.023497<br>\n3    0000_left_neural_foraminal_narrowing_l4_l5  0.510915    0.391577    0.092108<br>\n4    0000_left_neural_foraminal_narrowing_l5_s1  0.557290    0.308225    0.127156</p>\n<p>Solution<br>\nrow_id    normal_mild moderate    severe<br>\n0    0000_left_neural_foraminal_narrowing_l1_l2  1.0 0.0 0.0<br>\n1    0000_left_neural_foraminal_narrowing_l2_l3  0.0 1.0 0.0<br>\n2    0000_left_neural_foraminal_narrowing_l3_l4  0.0 1.0 0.0<br>\n3    0000_left_neural_foraminal_narrowing_l4_l5  0.0 0.0 1.0<br>\n4    0000_left_neural_foraminal_narrowing_l5_s1  0.0 1.0 0.0</p>",
      "rawMarkdown": "I made successful submissions, I pass this to latest scoring funciton, and it's looking for a mysterious column\n\nsample_weight=solution.loc[condition_indices, 'sample_weight'].values\n\nneither the Submission or the Solution supposed to have a sample_weight?\n\nHow is this calculated?\n\nSubmission\nrow_id\tnormal_mild\tmoderate\tsevere\n0\t0000_left_neural_foraminal_narrowing_l1_l2\t0.958009\t0.040073\t0.000177\n1\t0000_left_neural_foraminal_narrowing_l2_l3\t0.889687\t0.101157\t0.003351\n2\t0000_left_neural_foraminal_narrowing_l3_l4\t0.701213\t0.266386\t0.023497\n3\t0000_left_neural_foraminal_narrowing_l4_l5\t0.510915\t0.391577\t0.092108\n4\t0000_left_neural_foraminal_narrowing_l5_s1\t0.557290\t0.308225\t0.127156\n\nSolution\nrow_id\tnormal_mild\tmoderate\tsevere\n0\t0000_left_neural_foraminal_narrowing_l1_l2\t1.0\t0.0\t0.0\n1\t0000_left_neural_foraminal_narrowing_l2_l3\t0.0\t1.0\t0.0\n2\t0000_left_neural_foraminal_narrowing_l3_l4\t0.0\t1.0\t0.0\n3\t0000_left_neural_foraminal_narrowing_l4_l5\t0.0\t0.0\t1.0\n4\t0000_left_neural_foraminal_narrowing_l5_s1\t0.0\t1.0\t0.0",
      "votes": null
    },
    {
      "id": "2878738",
      "postDate": "06/19/2024 06:10:49",
      "content": "<p>You can see the full details in <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/overview/evaluation\" target=\"_blank\">Evaluation</a>, but to your question:</p>\n<blockquote>\n  <p>The sample weights are as follows:<br>\n  1 for normal/mild.<br>\n  2 for moderate.<br>\n  4 for severe.</p>\n</blockquote>\n<p>In the <code>score</code> function that the competition gave, you need to add a <code>sample_weight</code> column with values as above for each \"true\" label. Here is an example solution dataframe:</p>\n<table>\n<thead>\n<tr>\n<th>row_id</th>\n<th>normal_mild</th>\n<th>moderate</th>\n<th>severe</th>\n<th>sample_weight</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0000_left_neural_foraminal_narrowing_l1_l2</td>\n<td>1.0</td>\n<td>0.0</td>\n<td>0.0</td>\n<td>1</td>\n</tr>\n<tr>\n<td>0000_left_neural_foraminal_narrowing_l2_l3</td>\n<td>0.0</td>\n<td>1.0</td>\n<td>0.0</td>\n<td>2</td>\n</tr>\n<tr>\n<td>0000_left_neural_foraminal_narrowing_l3_l4</td>\n<td>0.0</td>\n<td>1.0</td>\n<td>0.0</td>\n<td>2</td>\n</tr>\n<tr>\n<td>0000_left_neural_foraminal_narrowing_l4_l5</td>\n<td>0.0</td>\n<td>0.0</td>\n<td>1.0</td>\n<td>4</td>\n</tr>\n<tr>\n<td>0000_left_neural_foraminal_narrowing_l5_s1</td>\n<td>0.0</td>\n<td>1.0</td>\n<td>0.0</td>\n<td>2</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "You can see the full details in [Evaluation](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/overview/evaluation), but to your question:\n\n>The sample weights are as follows:\n1 for normal/mild.\n2 for moderate.\n4 for severe.\n\nIn the `score` function that the competition gave, you need to add a `sample_weight` column with values as above for each \"true\" label. Here is an example solution dataframe:\n|   row_id |   normal_mild |   moderate | severe | sample_weight |\n|-----------:|------------:|------------------:|----------------------:|--------:|\n| 0000_left_neural_foraminal_narrowing_l1_l2 | 1.0 | 0.0 | 0.0 | 1 |\n| 0000_left_neural_foraminal_narrowing_l2_l3 | 0.0 | 1.0 | 0.0 | 2 |\n| 0000_left_neural_foraminal_narrowing_l3_l4 | 0.0 | 1.0 | 0.0 | 2 |\n| 0000_left_neural_foraminal_narrowing_l4_l5 | 0.0 | 0.0 | 1.0 | 4 |\n| 0000_left_neural_foraminal_narrowing_l5_s1 | 0.0 | 1.0 | 0.0 | 2 |",
      "votes": null
    },
    {
      "id": "2879014",
      "postDate": "06/19/2024 10:22:21",
      "content": "<p>If you are confused on how to compute it here is a sample code that can be used for <a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791\" target=\"_blank\">Version 10</a>:</p>\n<pre><code> ():\n    target_cols = (sample_weights.keys())\n    pred = submission.copy() \n    \n    pred[target_cols] = pred[target_cols].div(pred[target_cols].(axis=), axis=)\n\n    \n    indexed_train_df = train_df.set_index(, verify_integrity=)\n\n    row_ids = pred[row_id_column_name]\n    study_ids = row_ids.apply( x: x.split()[])\n    locations = row_ids.apply( x: .join(x.split()[:]))\n\n    solution_data = np.zeros_like(pred[target_cols].values)\n    sample_weight_list = []\n    nan_row_ids = ()\n     idx, (row, study_id, location)  ((row_ids, study_ids, locations)):\n        severity = (indexed_train_df.at[(study_id), location]).replace(, ).lower()\n         severity  sample_weights:\n            solution_data[idx, target_cols.index(severity)] = \n            sample_weight_list.append(sample_weights[severity])\n        :\n            solution_data[idx] = np.nan\n            nan_row_ids.add(row)\n            sample_weight_list.append(np.nan)\n\n    solution = pd.DataFrame({\n        row_id_column_name: pred[row_id_column_name],\n        : sample_weight_list\n    })\n    solution[target_cols] = solution_data\n\n    \n    pred.loc[pred[row_id_column_name].isin(nan_row_ids), target_cols] = np.nan\n    \n    \n     score(solution.dropna().copy(), pred.dropna().copy(), row_id_column_name, any_severe_scalar)\n</code></pre>\n<p>This converts DataFrame from <code>train.csv</code> to the format expected by <code>score</code> function, removes row_ids that are not labelled (that is, are NaN in <code>train.csv</code>) and passes copies to <code>score</code> function (<a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791\" target=\"_blank\">Version 10</a>) which should be in copied into your scope by copying the code as is or importing it.</p>",
      "rawMarkdown": "If you are confused on how to compute it here is a sample code that can be used for [Version 10](https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791):\n```python\ndef score_from_train(\n    submission: pd.DataFrame,  # Pass submission.csv as a DataFrame\n    train_df: pd.DataFrame,  # Pass train.csv as a DataFrame\n    row_id_column_name=\"row_id\",\n    any_severe_scalar=1.0,\n    sample_weights: dict[str, int]={\"normal_mild\": 1, \"moderate\": 2, \"severe\": 4},\n):\n    target_cols = list(sample_weights.keys())\n    pred = submission.copy() # Copy to prevent changes in original\n    # Normalize values to have a sum of 1.0\n    pred[target_cols] = pred[target_cols].div(pred[target_cols].sum(axis=1), axis=0)\n\n    # Index the study_id in train_df\n    indexed_train_df = train_df.set_index(\"study_id\", verify_integrity=True)\n\n    row_ids = pred[row_id_column_name]\n    study_ids = row_ids.apply(lambda x: x.split('_')[0])\n    locations = row_ids.apply(lambda x: '_'.join(x.split('_')[1:]))\n\n    solution_data = np.zeros_like(pred[target_cols].values)\n    sample_weight_list = []\n    nan_row_ids = set()\n    for idx, (row, study_id, location) in enumerate(zip(row_ids, study_ids, locations)):\n        severity = str(indexed_train_df.at[int(study_id), location]).replace(\"/\", \"_\").lower()\n        if severity in sample_weights:\n            solution_data[idx, target_cols.index(severity)] = 1.0\n            sample_weight_list.append(sample_weights[severity])\n        else:\n            solution_data[idx] = np.nan\n            nan_row_ids.add(row)\n            sample_weight_list.append(np.nan)\n\n    solution = pd.DataFrame({\n        row_id_column_name: pred[row_id_column_name],\n        \"sample_weight\": sample_weight_list\n    })\n    solution[target_cols] = solution_data\n\n    # Change row_ids in nan_row_ids to np.nan\n    pred.loc[pred[row_id_column_name].isin(nan_row_ids), target_cols] = np.nan\n    # Remove nan rows and pass copy to score function\n    # score from https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791 (Version 10)\n    return score(solution.dropna().copy(), pred.dropna().copy(), row_id_column_name, any_severe_scalar)\n```\nThis converts DataFrame from `train.csv` to the format expected by `score` function, removes row_ids that are not labelled (that is, are NaN in `train.csv`) and passes copies to `score` function ([Version 10](https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791)) which should be in copied into your scope by copying the code as is or importing it.",
      "votes": null
    },
    {
      "id": "2879207",
      "postDate": "06/19/2024 12:33:15",
      "content": "<p>Thank you so much for all the help.</p>",
      "rawMarkdown": "Thank you so much for all the help.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2878738,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/19/2024 06:10:49",
      "content": "<p>You can see the full details in <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/overview/evaluation\" target=\"_blank\">Evaluation</a>, but to your question:</p>\n<blockquote>\n  <p>The sample weights are as follows:<br>\n  1 for normal/mild.<br>\n  2 for moderate.<br>\n  4 for severe.</p>\n</blockquote>\n<p>In the <code>score</code> function that the competition gave, you need to add a <code>sample_weight</code> column with values as above for each \"true\" label. Here is an example solution dataframe:</p>\n<table>\n<thead>\n<tr>\n<th>row_id</th>\n<th>normal_mild</th>\n<th>moderate</th>\n<th>severe</th>\n<th>sample_weight</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0000_left_neural_foraminal_narrowing_l1_l2</td>\n<td>1.0</td>\n<td>0.0</td>\n<td>0.0</td>\n<td>1</td>\n</tr>\n<tr>\n<td>0000_left_neural_foraminal_narrowing_l2_l3</td>\n<td>0.0</td>\n<td>1.0</td>\n<td>0.0</td>\n<td>2</td>\n</tr>\n<tr>\n<td>0000_left_neural_foraminal_narrowing_l3_l4</td>\n<td>0.0</td>\n<td>1.0</td>\n<td>0.0</td>\n<td>2</td>\n</tr>\n<tr>\n<td>0000_left_neural_foraminal_narrowing_l4_l5</td>\n<td>0.0</td>\n<td>0.0</td>\n<td>1.0</td>\n<td>4</td>\n</tr>\n<tr>\n<td>0000_left_neural_foraminal_narrowing_l5_s1</td>\n<td>0.0</td>\n<td>1.0</td>\n<td>0.0</td>\n<td>2</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 2879207,
          "author_name": "abualabed",
          "author_url": "",
          "post_date": "06/19/2024 12:33:15",
          "content": "<p>Thank you so much for all the help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2879014,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/19/2024 10:22:21",
      "content": "<p>If you are confused on how to compute it here is a sample code that can be used for <a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791\" target=\"_blank\">Version 10</a>:</p>\n<pre><code> ():\n    target_cols = (sample_weights.keys())\n    pred = submission.copy() \n    \n    pred[target_cols] = pred[target_cols].div(pred[target_cols].(axis=), axis=)\n\n    \n    indexed_train_df = train_df.set_index(, verify_integrity=)\n\n    row_ids = pred[row_id_column_name]\n    study_ids = row_ids.apply( x: x.split()[])\n    locations = row_ids.apply( x: .join(x.split()[:]))\n\n    solution_data = np.zeros_like(pred[target_cols].values)\n    sample_weight_list = []\n    nan_row_ids = ()\n     idx, (row, study_id, location)  ((row_ids, study_ids, locations)):\n        severity = (indexed_train_df.at[(study_id), location]).replace(, ).lower()\n         severity  sample_weights:\n            solution_data[idx, target_cols.index(severity)] = \n            sample_weight_list.append(sample_weights[severity])\n        :\n            solution_data[idx] = np.nan\n            nan_row_ids.add(row)\n            sample_weight_list.append(np.nan)\n\n    solution = pd.DataFrame({\n        row_id_column_name: pred[row_id_column_name],\n        : sample_weight_list\n    })\n    solution[target_cols] = solution_data\n\n    \n    pred.loc[pred[row_id_column_name].isin(nan_row_ids), target_cols] = np.nan\n    \n    \n     score(solution.dropna().copy(), pred.dropna().copy(), row_id_column_name, any_severe_scalar)\n</code></pre>\n<p>This converts DataFrame from <code>train.csv</code> to the format expected by <code>score</code> function, removes row_ids that are not labelled (that is, are NaN in <code>train.csv</code>) and passes copies to <code>score</code> function (<a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791\" target=\"_blank\">Version 10</a>) which should be in copied into your scope by copying the code as is or importing it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2878545": "I made successful submissions, I pass this to latest scoring funciton, and it's looking for a mysterious column\n\nsample_weight=solution.loc[condition_indices, 'sample_weight'].values\n\nneither the Submission or the Solution supposed to have a sample_weight?\n\nHow is this calculated?\n\nSubmission\nrow_id\tnormal_mild\tmoderate\tsevere\n0\t0000_left_neural_foraminal_narrowing_l1_l2\t0.958009\t0.040073\t0.000177\n1\t0000_left_neural_foraminal_narrowing_l2_l3\t0.889687\t0.101157\t0.003351\n2\t0000_left_neural_foraminal_narrowing_l3_l4\t0.701213\t0.266386\t0.023497\n3\t0000_left_neural_foraminal_narrowing_l4_l5\t0.510915\t0.391577\t0.092108\n4\t0000_left_neural_foraminal_narrowing_l5_s1\t0.557290\t0.308225\t0.127156\n\nSolution\nrow_id\tnormal_mild\tmoderate\tsevere\n0\t0000_left_neural_foraminal_narrowing_l1_l2\t1.0\t0.0\t0.0\n1\t0000_left_neural_foraminal_narrowing_l2_l3\t0.0\t1.0\t0.0\n2\t0000_left_neural_foraminal_narrowing_l3_l4\t0.0\t1.0\t0.0\n3\t0000_left_neural_foraminal_narrowing_l4_l5\t0.0\t0.0\t1.0\n4\t0000_left_neural_foraminal_narrowing_l5_s1\t0.0\t1.0\t0.0",
    "2878738": "You can see the full details in [Evaluation](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/overview/evaluation), but to your question:\n\n>The sample weights are as follows:\n1 for normal/mild.\n2 for moderate.\n4 for severe.\n\nIn the `score` function that the competition gave, you need to add a `sample_weight` column with values as above for each \"true\" label. Here is an example solution dataframe:\n|   row_id |   normal_mild |   moderate | severe | sample_weight |\n|-----------:|------------:|------------------:|----------------------:|--------:|\n| 0000_left_neural_foraminal_narrowing_l1_l2 | 1.0 | 0.0 | 0.0 | 1 |\n| 0000_left_neural_foraminal_narrowing_l2_l3 | 0.0 | 1.0 | 0.0 | 2 |\n| 0000_left_neural_foraminal_narrowing_l3_l4 | 0.0 | 1.0 | 0.0 | 2 |\n| 0000_left_neural_foraminal_narrowing_l4_l5 | 0.0 | 0.0 | 1.0 | 4 |\n| 0000_left_neural_foraminal_narrowing_l5_s1 | 0.0 | 1.0 | 0.0 | 2 |",
    "2879014": "If you are confused on how to compute it here is a sample code that can be used for [Version 10](https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791):\n```python\ndef score_from_train(\n    submission: pd.DataFrame,  # Pass submission.csv as a DataFrame\n    train_df: pd.DataFrame,  # Pass train.csv as a DataFrame\n    row_id_column_name=\"row_id\",\n    any_severe_scalar=1.0,\n    sample_weights: dict[str, int]={\"normal_mild\": 1, \"moderate\": 2, \"severe\": 4},\n):\n    target_cols = list(sample_weights.keys())\n    pred = submission.copy() # Copy to prevent changes in original\n    # Normalize values to have a sum of 1.0\n    pred[target_cols] = pred[target_cols].div(pred[target_cols].sum(axis=1), axis=0)\n\n    # Index the study_id in train_df\n    indexed_train_df = train_df.set_index(\"study_id\", verify_integrity=True)\n\n    row_ids = pred[row_id_column_name]\n    study_ids = row_ids.apply(lambda x: x.split('_')[0])\n    locations = row_ids.apply(lambda x: '_'.join(x.split('_')[1:]))\n\n    solution_data = np.zeros_like(pred[target_cols].values)\n    sample_weight_list = []\n    nan_row_ids = set()\n    for idx, (row, study_id, location) in enumerate(zip(row_ids, study_ids, locations)):\n        severity = str(indexed_train_df.at[int(study_id), location]).replace(\"/\", \"_\").lower()\n        if severity in sample_weights:\n            solution_data[idx, target_cols.index(severity)] = 1.0\n            sample_weight_list.append(sample_weights[severity])\n        else:\n            solution_data[idx] = np.nan\n            nan_row_ids.add(row)\n            sample_weight_list.append(np.nan)\n\n    solution = pd.DataFrame({\n        row_id_column_name: pred[row_id_column_name],\n        \"sample_weight\": sample_weight_list\n    })\n    solution[target_cols] = solution_data\n\n    # Change row_ids in nan_row_ids to np.nan\n    pred.loc[pred[row_id_column_name].isin(nan_row_ids), target_cols] = np.nan\n    # Remove nan rows and pass copy to score function\n    # score from https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791 (Version 10)\n    return score(solution.dropna().copy(), pred.dropna().copy(), row_id_column_name, any_severe_scalar)\n```\nThis converts DataFrame from `train.csv` to the format expected by `score` function, removes row_ids that are not labelled (that is, are NaN in `train.csv`) and passes copies to `score` function ([Version 10](https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791)) which should be in copied into your scope by copying the code as is or importing it.",
    "2879207": "Thank you so much for all the help."
  },
  "source": "meta"
}