{
  "id": 507498,
  "title": " Is any_severe_scalar public?",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/507498",
  "author_name": "",
  "post_date": "2024-05-26T05:39:35.048856100Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/abhinavsuri\" target=\"_blank\">@abhinavsuri</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nThank you for publishing such a great contest!　<br>\nI would like to calculate metrics locally, so I would like to know if any_severe_scalar is public.</p>",
  "messages": [
    {
      "id": "2836761",
      "postDate": "05/26/2024 05:39:35",
      "content": "<p><a href=\"https://www.kaggle.com/abhinavsuri\" target=\"_blank\">@abhinavsuri</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nThank you for publishing such a great contest!　<br>\nI would like to calculate metrics locally, so I would like to know if any_severe_scalar is public.</p>",
      "rawMarkdown": "abhinavsuri @sohier \nThank you for publishing such a great contest!　\nI would like to calculate metrics locally, so I would like to know if any_severe_scalar is public.",
      "votes": null
    },
    {
      "id": "2841552",
      "postDate": "05/28/2024 15:58:40",
      "content": "<p>It's just <code>1.0</code>. Thanks for asking, I've added a note to that effect to the evaluation tab as well.</p>",
      "rawMarkdown": "It's just `1.0`. Thanks for asking, I've added a note to that effect to the evaluation tab as well.",
      "votes": null
    },
    {
      "id": "2866258",
      "postDate": "06/11/2024 08:22:51",
      "content": "<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> How can you get solutions.csv locally? If you create it to test your code on a part of train_data with the help of train.csv, please guide me on how to do the same. Thank you. You r response will be invaluable.</p>",
      "rawMarkdown": "abebe9849 How can you get solutions.csv locally? If you create it to test your code on a part of train_data with the help of train.csv, please guide me on how to do the same. Thank you. You r response will be invaluable.",
      "votes": null
    },
    {
      "id": "2866301",
      "postDate": "06/11/2024 08:38:27",
      "content": "<pre><code> metric  CALC_score \n ():\n    new_df = pd.DataFrame()\n    tra_df = (pd.read_csv().columns[:])\n    col = []\n    c_ = []\n     i  test_stusy:\n         j  tra_df:\n            col.append()\n            c_.append()\n\n    new_df[]=col\n    new_df[]=c_\n    new_df[] = new_df[].astype()\n    new_df[]=\n    new_df[]=\n    new_df[]=\n    new_df[[,,]]=test_preds.reshape(-,)\n    new_df[]=folds_tmp.iloc[:,:-].to_numpy().reshape((-,))\n    new_df = new_df[new_df[]!=-].reset_index(drop=)\n    GT = new_df.iloc[:,[,-]].copy()\n    GT[[,,]]=np.eye()[GT[].to_numpy().astype(np.uint8)]\n    GT[]=**GT[].to_numpy()\n\n    GT = GT.iloc[:,[,,,,]]\n    metirc_ = CALC_score(GT,new_df.iloc[:,[,,,]],row_id_column_name=)\n     metirc_,new_df\n\nval_folds = val_folds.groupby().first().reset_index() \n\nval_folds = val_folds.reset_index(drop=)\nc_,val_df_ = make_calc(val_folds[].unique(),preds,val_folds)\n\n</code></pre>\n<p>I am using this to calculate local metirc</p>",
      "rawMarkdown": "```python\nfrom metric import CALC_score # same as https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549\ndef make_calc(test_stusy,test_preds,folds_tmp):\n    new_df = pd.DataFrame()\n    tra_df = list(pd.read_csv(\"/home/u094724e/rsna2024/data/train.csv\").columns[1:])\n    col = []\n    c_ = []\n    for i in test_stusy:\n        for j in tra_df:\n            col.append(f\"{i}_{j}\")\n            c_.append(f\"{i}\")\n            \n    new_df[\"row_id\"]=col\n    new_df[\"study_id\"]=c_\n    new_df[\"row_id\"] = new_df[\"row_id\"].astype(\"str\")\n    new_df[\"normal_mild\"]=0\n    new_df[\"moderate\"]=0\n    new_df[\"severe\"]=0\n    new_df[[\"normal_mild\",\"moderate\",\"severe\"]]=test_preds.reshape(-1,3)\n    new_df[\"GT\"]=folds_tmp.iloc[:,2:-1].to_numpy().reshape((-1,))\n    new_df = new_df[new_df[\"GT\"]!=-1].reset_index(drop=True)\n    GT = new_df.iloc[:,[0,-1]].copy()\n    GT[[\"normal_mild\",\"moderate\",\"severe\"]]=np.eye(3)[GT[\"GT\"].to_numpy().astype(np.uint8)]\n    GT[\"sample_weight\"]=2**GT[\"GT\"].to_numpy()\n\n    GT = GT.iloc[:,[0,2,3,4,5]]\n    metirc_ = CALC_score(GT,new_df.iloc[:,[0,2,3,4]],row_id_column_name=\"row_id\")\n    return metirc_,new_df\n\nval_folds = val_folds.groupby(\"study_id\").first().reset_index() \n\nval_folds = val_folds.reset_index(drop=True)\nc_,val_df_ = make_calc(val_folds[\"study_id\"].unique(),preds,val_folds)\n#preds.... num_study_id,75cls(5*5*3)\n```\n\nI am using this to calculate local metirc",
      "votes": null
    },
    {
      "id": "2867026",
      "postDate": "06/11/2024 15:47:34",
      "content": "<p>I thank you a lot. Although it is daunting to me, it will be helpful to me in the upcoming days. Thank you again for performing greatly but still being humble.</p>",
      "rawMarkdown": "I thank you a lot. Although it is daunting to me, it will be helpful to me in the upcoming days. Thank you again for performing greatly but still being humble.",
      "votes": null
    },
    {
      "id": "2879004",
      "postDate": "06/19/2024 10:19:11",
      "content": "<p>Here is an updated code that can be used for Version 10:</p>\n<pre><code> ():\n    target_cols = (sample_weights.keys())\n    pred = submission.copy() \n    \n    pred[target_cols] = pred[target_cols].div(pred[target_cols].(axis=), axis=)\n\n    \n    indexed_train_df = train_df.set_index(, verify_integrity=)\n\n    row_ids = pred[row_id_column_name]\n    study_ids = row_ids.apply( x: x.split()[])\n    locations = row_ids.apply( x: .join(x.split()[:]))\n\n    solution_data = np.zeros_like(pred[target_cols].values)\n    sample_weight_list = []\n    nan_row_ids = ()\n     idx, (row, study_id, location)  ((row_ids, study_ids, locations)):\n        severity = (indexed_train_df.at[(study_id), location]).replace(, ).lower()\n         severity  sample_weights:\n            solution_data[idx, target_cols.index(severity)] = \n            sample_weight_list.append(sample_weights[severity])\n        :\n            solution_data[idx] = np.nan\n            nan_row_ids.add(row)\n            sample_weight_list.append(np.nan)\n\n    solution = pd.DataFrame({\n        row_id_column_name: pred[row_id_column_name],\n        : sample_weight_list\n    })\n    solution[target_cols] = solution_data\n\n    \n    pred.loc[pred[row_id_column_name].isin(nan_row_ids), target_cols] = np.nan\n    \n    \n     score(solution.dropna().copy(), pred.dropna().copy(), row_id_column_name, any_severe_scalar)\n</code></pre>\n<p>This converts DataFrame from <code>train.csv</code> to the format expected by <code>score</code> function, removes row_ids that are not labelled (that is, are NaN in <code>train.csv</code>) and passes copies to <code>score</code> function (<a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791\" target=\"_blank\">Version 10</a>)</p>",
      "rawMarkdown": "Here is an updated code that can be used for Version 10:\n```python\ndef score_from_train(\n    submission: pd.DataFrame,  # Pass submission.csv as a DataFrame\n    train_df: pd.DataFrame,  # Pass train.csv as a DataFrame\n    row_id_column_name=\"row_id\",\n    any_severe_scalar=1.0,\n    sample_weights: dict[str, int]={\"normal_mild\": 1, \"moderate\": 2, \"severe\": 4},\n):\n    target_cols = list(sample_weights.keys())\n    pred = submission.copy() # Copy to prevent changes in original\n    # Normalize values to have a sum of 1.0\n    pred[target_cols] = pred[target_cols].div(pred[target_cols].sum(axis=1), axis=0)\n\n    # Index the study_id in train_df\n    indexed_train_df = train_df.set_index(\"study_id\", verify_integrity=True)\n\n    row_ids = pred[row_id_column_name]\n    study_ids = row_ids.apply(lambda x: x.split('_')[0])\n    locations = row_ids.apply(lambda x: '_'.join(x.split('_')[1:]))\n\n    solution_data = np.zeros_like(pred[target_cols].values)\n    sample_weight_list = []\n    nan_row_ids = set()\n    for idx, (row, study_id, location) in enumerate(zip(row_ids, study_ids, locations)):\n        severity = str(indexed_train_df.at[int(study_id), location]).replace(\"/\", \"_\").lower()\n        if severity in sample_weights:\n            solution_data[idx, target_cols.index(severity)] = 1.0\n            sample_weight_list.append(sample_weights[severity])\n        else:\n            solution_data[idx] = np.nan\n            nan_row_ids.add(row)\n            sample_weight_list.append(np.nan)\n\n    solution = pd.DataFrame({\n        row_id_column_name: pred[row_id_column_name],\n        \"sample_weight\": sample_weight_list\n    })\n    solution[target_cols] = solution_data\n\n    # Change row_ids in nan_row_ids to np.nan\n    pred.loc[pred[row_id_column_name].isin(nan_row_ids), target_cols] = np.nan\n    # Remove nan rows and pass copy to score function\n    # score from https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791 (Version 10)\n    return score(solution.dropna().copy(), pred.dropna().copy(), row_id_column_name, any_severe_scalar)\n```\nThis converts DataFrame from `train.csv` to the format expected by `score` function, removes row_ids that are not labelled (that is, are NaN in `train.csv`) and passes copies to `score` function ([Version 10](https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791))",
      "votes": null
    },
    {
      "id": "2906198",
      "postDate": "07/05/2024 12:33:39",
      "content": "<p>Here is a notebook simulating the metric locally: <a href=\"https://www.kaggle.com/code/sadidul012/rsna24-lsdc-validation-scoring-from-trainset\" target=\"_blank\">https://www.kaggle.com/code/sadidul012/rsna24-lsdc-validation-scoring-from-trainset</a></p>",
      "rawMarkdown": "Here is a notebook simulating the metric locally: https://www.kaggle.com/code/sadidul012/rsna24-lsdc-validation-scoring-from-trainset",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2841552,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "05/28/2024 15:58:40",
      "content": "<p>It's just <code>1.0</code>. Thanks for asking, I've added a note to that effect to the evaluation tab as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2866258,
      "author_name": "devsya",
      "author_url": "",
      "post_date": "06/11/2024 08:22:51",
      "content": "<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> How can you get solutions.csv locally? If you create it to test your code on a part of train_data with the help of train.csv, please guide me on how to do the same. Thank you. You r response will be invaluable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2866301,
          "author_name": "abebe9849",
          "author_url": "",
          "post_date": "06/11/2024 08:38:27",
          "content": "<pre><code> metric  CALC_score \n ():\n    new_df = pd.DataFrame()\n    tra_df = (pd.read_csv().columns[:])\n    col = []\n    c_ = []\n     i  test_stusy:\n         j  tra_df:\n            col.append()\n            c_.append()\n\n    new_df[]=col\n    new_df[]=c_\n    new_df[] = new_df[].astype()\n    new_df[]=\n    new_df[]=\n    new_df[]=\n    new_df[[,,]]=test_preds.reshape(-,)\n    new_df[]=folds_tmp.iloc[:,:-].to_numpy().reshape((-,))\n    new_df = new_df[new_df[]!=-].reset_index(drop=)\n    GT = new_df.iloc[:,[,-]].copy()\n    GT[[,,]]=np.eye()[GT[].to_numpy().astype(np.uint8)]\n    GT[]=**GT[].to_numpy()\n\n    GT = GT.iloc[:,[,,,,]]\n    metirc_ = CALC_score(GT,new_df.iloc[:,[,,,]],row_id_column_name=)\n     metirc_,new_df\n\nval_folds = val_folds.groupby().first().reset_index() \n\nval_folds = val_folds.reset_index(drop=)\nc_,val_df_ = make_calc(val_folds[].unique(),preds,val_folds)\n\n</code></pre>\n<p>I am using this to calculate local metirc</p>",
          "votes": null,
          "replies": [
            {
              "id": 2867026,
              "author_name": "devsya",
              "author_url": "",
              "post_date": "06/11/2024 15:47:34",
              "content": "<p>I thank you a lot. Although it is daunting to me, it will be helpful to me in the upcoming days. Thank you again for performing greatly but still being humble.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2879004,
          "author_name": "coderrkj",
          "author_url": "",
          "post_date": "06/19/2024 10:19:11",
          "content": "<p>Here is an updated code that can be used for Version 10:</p>\n<pre><code> ():\n    target_cols = (sample_weights.keys())\n    pred = submission.copy() \n    \n    pred[target_cols] = pred[target_cols].div(pred[target_cols].(axis=), axis=)\n\n    \n    indexed_train_df = train_df.set_index(, verify_integrity=)\n\n    row_ids = pred[row_id_column_name]\n    study_ids = row_ids.apply( x: x.split()[])\n    locations = row_ids.apply( x: .join(x.split()[:]))\n\n    solution_data = np.zeros_like(pred[target_cols].values)\n    sample_weight_list = []\n    nan_row_ids = ()\n     idx, (row, study_id, location)  ((row_ids, study_ids, locations)):\n        severity = (indexed_train_df.at[(study_id), location]).replace(, ).lower()\n         severity  sample_weights:\n            solution_data[idx, target_cols.index(severity)] = \n            sample_weight_list.append(sample_weights[severity])\n        :\n            solution_data[idx] = np.nan\n            nan_row_ids.add(row)\n            sample_weight_list.append(np.nan)\n\n    solution = pd.DataFrame({\n        row_id_column_name: pred[row_id_column_name],\n        : sample_weight_list\n    })\n    solution[target_cols] = solution_data\n\n    \n    pred.loc[pred[row_id_column_name].isin(nan_row_ids), target_cols] = np.nan\n    \n    \n     score(solution.dropna().copy(), pred.dropna().copy(), row_id_column_name, any_severe_scalar)\n</code></pre>\n<p>This converts DataFrame from <code>train.csv</code> to the format expected by <code>score</code> function, removes row_ids that are not labelled (that is, are NaN in <code>train.csv</code>) and passes copies to <code>score</code> function (<a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791\" target=\"_blank\">Version 10</a>)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2906198,
      "author_name": "sadidul012",
      "author_url": "",
      "post_date": "07/05/2024 12:33:39",
      "content": "<p>Here is a notebook simulating the metric locally: <a href=\"https://www.kaggle.com/code/sadidul012/rsna24-lsdc-validation-scoring-from-trainset\" target=\"_blank\">https://www.kaggle.com/code/sadidul012/rsna24-lsdc-validation-scoring-from-trainset</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2836761": "abhinavsuri @sohier \nThank you for publishing such a great contest!　\nI would like to calculate metrics locally, so I would like to know if any_severe_scalar is public.",
    "2841552": "It's just `1.0`. Thanks for asking, I've added a note to that effect to the evaluation tab as well.",
    "2866258": "abebe9849 How can you get solutions.csv locally? If you create it to test your code on a part of train_data with the help of train.csv, please guide me on how to do the same. Thank you. You r response will be invaluable.",
    "2866301": "```python\nfrom metric import CALC_score # same as https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549\ndef make_calc(test_stusy,test_preds,folds_tmp):\n    new_df = pd.DataFrame()\n    tra_df = list(pd.read_csv(\"/home/u094724e/rsna2024/data/train.csv\").columns[1:])\n    col = []\n    c_ = []\n    for i in test_stusy:\n        for j in tra_df:\n            col.append(f\"{i}_{j}\")\n            c_.append(f\"{i}\")\n            \n    new_df[\"row_id\"]=col\n    new_df[\"study_id\"]=c_\n    new_df[\"row_id\"] = new_df[\"row_id\"].astype(\"str\")\n    new_df[\"normal_mild\"]=0\n    new_df[\"moderate\"]=0\n    new_df[\"severe\"]=0\n    new_df[[\"normal_mild\",\"moderate\",\"severe\"]]=test_preds.reshape(-1,3)\n    new_df[\"GT\"]=folds_tmp.iloc[:,2:-1].to_numpy().reshape((-1,))\n    new_df = new_df[new_df[\"GT\"]!=-1].reset_index(drop=True)\n    GT = new_df.iloc[:,[0,-1]].copy()\n    GT[[\"normal_mild\",\"moderate\",\"severe\"]]=np.eye(3)[GT[\"GT\"].to_numpy().astype(np.uint8)]\n    GT[\"sample_weight\"]=2**GT[\"GT\"].to_numpy()\n\n    GT = GT.iloc[:,[0,2,3,4,5]]\n    metirc_ = CALC_score(GT,new_df.iloc[:,[0,2,3,4]],row_id_column_name=\"row_id\")\n    return metirc_,new_df\n\nval_folds = val_folds.groupby(\"study_id\").first().reset_index() \n\nval_folds = val_folds.reset_index(drop=True)\nc_,val_df_ = make_calc(val_folds[\"study_id\"].unique(),preds,val_folds)\n#preds.... num_study_id,75cls(5*5*3)\n```\n\nI am using this to calculate local metirc",
    "2867026": "I thank you a lot. Although it is daunting to me, it will be helpful to me in the upcoming days. Thank you again for performing greatly but still being humble.",
    "2879004": "Here is an updated code that can be used for Version 10:\n```python\ndef score_from_train(\n    submission: pd.DataFrame,  # Pass submission.csv as a DataFrame\n    train_df: pd.DataFrame,  # Pass train.csv as a DataFrame\n    row_id_column_name=\"row_id\",\n    any_severe_scalar=1.0,\n    sample_weights: dict[str, int]={\"normal_mild\": 1, \"moderate\": 2, \"severe\": 4},\n):\n    target_cols = list(sample_weights.keys())\n    pred = submission.copy() # Copy to prevent changes in original\n    # Normalize values to have a sum of 1.0\n    pred[target_cols] = pred[target_cols].div(pred[target_cols].sum(axis=1), axis=0)\n\n    # Index the study_id in train_df\n    indexed_train_df = train_df.set_index(\"study_id\", verify_integrity=True)\n\n    row_ids = pred[row_id_column_name]\n    study_ids = row_ids.apply(lambda x: x.split('_')[0])\n    locations = row_ids.apply(lambda x: '_'.join(x.split('_')[1:]))\n\n    solution_data = np.zeros_like(pred[target_cols].values)\n    sample_weight_list = []\n    nan_row_ids = set()\n    for idx, (row, study_id, location) in enumerate(zip(row_ids, study_ids, locations)):\n        severity = str(indexed_train_df.at[int(study_id), location]).replace(\"/\", \"_\").lower()\n        if severity in sample_weights:\n            solution_data[idx, target_cols.index(severity)] = 1.0\n            sample_weight_list.append(sample_weights[severity])\n        else:\n            solution_data[idx] = np.nan\n            nan_row_ids.add(row)\n            sample_weight_list.append(np.nan)\n\n    solution = pd.DataFrame({\n        row_id_column_name: pred[row_id_column_name],\n        \"sample_weight\": sample_weight_list\n    })\n    solution[target_cols] = solution_data\n\n    # Change row_ids in nan_row_ids to np.nan\n    pred.loc[pred[row_id_column_name].isin(nan_row_ids), target_cols] = np.nan\n    # Remove nan rows and pass copy to score function\n    # score from https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791 (Version 10)\n    return score(solution.dropna().copy(), pred.dropna().copy(), row_id_column_name, any_severe_scalar)\n```\nThis converts DataFrame from `train.csv` to the format expected by `score` function, removes row_ids that are not labelled (that is, are NaN in `train.csv`) and passes copies to `score` function ([Version 10](https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791))",
    "2906198": "Here is a notebook simulating the metric locally: https://www.kaggle.com/code/sadidul012/rsna24-lsdc-validation-scoring-from-trainset"
  },
  "source": "meta"
}