{
  "id": 521786,
  "title": "How calculate metrics before submit solution?",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/521786",
  "author_name": "",
  "post_date": "2024-07-22T22:08:16.992588100Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I did run my notebook (and I can see my public score) so everything works fine. </p>\n<p>However I'm running my notebook in Visual Studio (locally) and I would like to calculate my public score before submission and double check if there is a significan improvement.</p>\n<p>I found this <a href=\"https://www.kaggle.com/code/sadidul012/rsna24-lsdc-validation-scoring-from-trainset\" target=\"_blank\">https://www.kaggle.com/code/sadidul012/rsna24-lsdc-validation-scoring-from-trainset</a> as something close to what I want (except submission is random and not predicted). Any clue?</p>",
  "messages": [
    {
      "id": "2932386",
      "postDate": "07/22/2024 22:08:16",
      "content": "<p>I did run my notebook (and I can see my public score) so everything works fine. </p>\n<p>However I'm running my notebook in Visual Studio (locally) and I would like to calculate my public score before submission and double check if there is a significan improvement.</p>\n<p>I found this <a href=\"https://www.kaggle.com/code/sadidul012/rsna24-lsdc-validation-scoring-from-trainset\" target=\"_blank\">https://www.kaggle.com/code/sadidul012/rsna24-lsdc-validation-scoring-from-trainset</a> as something close to what I want (except submission is random and not predicted). Any clue?</p>",
      "rawMarkdown": "I did run my notebook (and I can see my public score) so everything works fine. \n\nHowever I'm running my notebook in Visual Studio (locally) and I would like to calculate my public score before submission and double check if there is a significan improvement.\n\nI found this https://www.kaggle.com/code/sadidul012/rsna24-lsdc-validation-scoring-from-trainset as something close to what I want (except submission is random and not predicted). Any clue?",
      "votes": null
    },
    {
      "id": "2932414",
      "postDate": "07/22/2024 23:06:12",
      "content": "<p>Predict and introduce your \"submission\" (predicted scores) against the \"solution\" (target labels). The weights are 1,2 and 4.<br>\n<code>random_pred = your_prediction</code><br>\nActually random_pred is not random there, is just naive.</p>",
      "rawMarkdown": "Predict and introduce your \"submission\" (predicted scores) against the \"solution\" (target labels). The weights are 1,2 and 4.\n`random_pred = your_prediction`\nActually random_pred is not random there, is just naive.",
      "votes": null
    },
    {
      "id": "2932540",
      "postDate": "07/23/2024 05:11:17",
      "content": "<p>If your question is how to use the [metric given by the competition host](from <a href=\"https://www.kaggle.com/datasets/seoyunje/kaggle-rsna-2024-metric\" target=\"_blank\">https://www.kaggle.com/datasets/seoyunje/kaggle-rsna-2024-metric</a>) to see the CV score for a test run on train images, then use the following (first part is copied from the metric notebook linked above):</p>\n<pre><code> numpy  np\n pandas  pd\n pandas.api.types\n sklearn.metrics\n\n\n ():\n    \n\n\n () -&gt; :\n    \n     injury_condition  [, , ]:\n         injury_condition  full_location:\n             injury_condition\n     ValueError()\n\n\n () -&gt; :\n    \n\n    target_levels = [, , ]\n\n    \n      pandas.api.types.is_numeric_dtype(submission[target_levels].values):\n         ParticipantVisibleError()\n\n      np.isfinite(submission[target_levels].values).():\n         ParticipantVisibleError()\n\n     solution[target_levels].().() &lt; :\n         ParticipantVisibleError()\n     submission[target_levels].().() &lt; :\n         ParticipantVisibleError()\n\n    solution[] = solution[].apply( x: x.split()[])\n    solution[] = solution[].apply( x: .join(x.split()[:]))\n    solution[] = solution[].apply(get_condition)\n\n     solution[row_id_column_name]\n     submission[row_id_column_name]\n     (submission.columns) == (target_levels)\n\n    submission[] = solution[]\n    submission[] = solution[]\n    submission[] = solution[]\n\n    condition_losses = []\n    condition_weights = []\n     condition  [, , ]:\n        condition_indices = solution.loc[solution[] == condition].index.values\n        condition_loss = sklearn.metrics.log_loss(\n            y_true=solution.loc[condition_indices, target_levels].values,\n            y_pred=submission.loc[condition_indices, target_levels].values,\n            sample_weight=solution.loc[condition_indices, ].values\n        )\n        condition_losses.append(condition_loss)\n        condition_weights.append()\n\n    any_severe_spinal_labels = pd.Series(solution.loc[solution[] == ].groupby()[].())\n    any_severe_spinal_weights = pd.Series(solution.loc[solution[] == ].groupby()[].())\n    any_severe_spinal_predictions = pd.Series(submission.loc[submission[] == ].groupby()[].())\n    any_severe_spinal_loss = sklearn.metrics.log_loss(\n        y_true=any_severe_spinal_labels,\n        y_pred=any_severe_spinal_predictions,\n        sample_weight=any_severe_spinal_weights\n    )\n    condition_losses.append(any_severe_spinal_loss)\n    condition_weights.append(any_severe_scalar)\n     np.average(condition_losses, weights=condition_weights)\n\n\n ():\n    \n    target_cols = (sample_weights.keys())\n    pred = submission.copy() \n    \n    pred[target_cols] = pred[target_cols].div(pred[target_cols].(axis=), axis=)\n\n    \n    indexed_train_df = train_df.set_index(, verify_integrity=)\n\n    row_ids = pred[row_id_column_name]\n    study_ids = row_ids.apply( x: x.split()[])\n    locations = row_ids.apply( x: .join(x.split()[:]))\n\n    solution_data = np.zeros_like(pred[target_cols].values)\n    sample_weight_list = []\n    nan_row_ids = ()\n     idx, (row, study_id, location)  ((row_ids, study_ids, locations)):\n        severity = (indexed_train_df.at[(study_id), location]).replace(, ).lower()\n         severity  sample_weights:\n            solution_data[idx, target_cols.index(severity)] = \n            sample_weight_list.append(sample_weights[severity])\n        :\n            solution_data[idx] = np.nan\n            nan_row_ids.add(row)\n            sample_weight_list.append(np.nan)\n\n    solution = pd.DataFrame({\n        row_id_column_name: pred[row_id_column_name],\n        : sample_weight_list\n    })\n    solution[target_cols] = solution_data\n\n    \n    pred.loc[pred[row_id_column_name].isin(nan_row_ids), target_cols] = np.nan\n    \n    \n     score(solution.dropna().copy(), pred.dropna().copy(), row_id_column_name, any_severe_scalar)\n</code></pre>\n<p>Create a flag for testing on train images: <code>FAKE_TEST = len(sample_sub) &lt;= 25</code> this will ensure that during submission of inference notebook the actual test data is used.</p>\n<p>In your notebook swap path to images and other test data dataframes to train path and data. Example:</p>\n<pre><code> FAKE_TEST:\n    df = pd.read_csv(DATA_PATH / )\n    study_ids, base_path = (df[].unique()), DATA_PATH / \n    (, (study_ids))\n</code></pre>\n<p>Finally, do the test after <code>submission.csv</code> is created:</p>\n<pre><code> FAKE_TEST:\n    \n    (score_from_train(pd.read_csv(), pd.read_csv(DATA_PATH / )))\n</code></pre>\n<p>Here are my public notebooks, where I have done this:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/coderrkj/lumbar-competition-debugging-notebook\" target=\"_blank\">https://www.kaggle.com/code/coderrkj/lumbar-competition-debugging-notebook</a></li>\n<li><a href=\"https://www.kaggle.com/coderrkj/rsna-pytorch-test-infer\" target=\"_blank\">https://www.kaggle.com/coderrkj/rsna-pytorch-test-infer</a></li>\n<li><a href=\"https://www.kaggle.com/coderrkj/rsna2024-lsdc-submission-baseline\" target=\"_blank\">https://www.kaggle.com/coderrkj/rsna2024-lsdc-submission-baseline</a></li>\n</ul>",
      "rawMarkdown": "If your question is how to use the [metric given by the competition host](from https://www.kaggle.com/datasets/seoyunje/kaggle-rsna-2024-metric) to see the CV score for a test run on train images, then use the following (first part is copied from the metric notebook linked above):\n\n```python\nimport numpy as np\nimport pandas as pd\nimport pandas.api.types\nimport sklearn.metrics\n\n\nclass ParticipantVisibleError(Exception):\n    pass\n\n\ndef get_condition(full_location: str) -> str:\n    # Given an input like spinal_canal_stenosis_l1_l2 extracts 'spinal'\n    for injury_condition in ['spinal', 'foraminal', 'subarticular']:\n        if injury_condition in full_location:\n            return injury_condition\n    raise ValueError(f'condition not found in {full_location}')\n\n\ndef score(\n        solution: pd.DataFrame,\n        submission: pd.DataFrame,\n        row_id_column_name: str,\n        any_severe_scalar: float\n    ) -> float:\n    '''\n    Pseudocode:\n    1. Calculate the sample weighted log loss for each medical condition:\n    2. Derive a new any_severe label.\n    3. Calculate the sample weighted log loss for the new any_severe label.\n    4. Return the average of all of the label group log losses as the final score, normalized for the number of columns in each group.\n       This mitigates the impact of spinal stenosis having only half as many columns as the other two conditions.\n    '''\n\n    target_levels = ['normal_mild', 'moderate', 'severe']\n\n    # Run basic QC checks on the inputs\n    if not pandas.api.types.is_numeric_dtype(submission[target_levels].values):\n        raise ParticipantVisibleError('All submission values must be numeric')\n\n    if not np.isfinite(submission[target_levels].values).all():\n        raise ParticipantVisibleError('All submission values must be finite')\n\n    if solution[target_levels].min().min() < 0:\n        raise ParticipantVisibleError('All labels must be at least zero')\n    if submission[target_levels].min().min() < 0:\n        raise ParticipantVisibleError('All predictions must be at least zero')\n\n    solution['study_id'] = solution['row_id'].apply(lambda x: x.split('_')[0])\n    solution['location'] = solution['row_id'].apply(lambda x: '_'.join(x.split('_')[1:]))\n    solution['condition'] = solution['row_id'].apply(get_condition)\n\n    del solution[row_id_column_name]\n    del submission[row_id_column_name]\n    assert sorted(submission.columns) == sorted(target_levels)\n\n    submission['study_id'] = solution['study_id']\n    submission['location'] = solution['location']\n    submission['condition'] = solution['condition']\n\n    condition_losses = []\n    condition_weights = []\n    for condition in ['spinal', 'foraminal', 'subarticular']:\n        condition_indices = solution.loc[solution['condition'] == condition].index.values\n        condition_loss = sklearn.metrics.log_loss(\n            y_true=solution.loc[condition_indices, target_levels].values,\n            y_pred=submission.loc[condition_indices, target_levels].values,\n            sample_weight=solution.loc[condition_indices, 'sample_weight'].values\n        )\n        condition_losses.append(condition_loss)\n        condition_weights.append(1)\n\n    any_severe_spinal_labels = pd.Series(solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['severe'].max())\n    any_severe_spinal_weights = pd.Series(solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['sample_weight'].max())\n    any_severe_spinal_predictions = pd.Series(submission.loc[submission['condition'] == 'spinal'].groupby('study_id')['severe'].max())\n    any_severe_spinal_loss = sklearn.metrics.log_loss(\n        y_true=any_severe_spinal_labels,\n        y_pred=any_severe_spinal_predictions,\n        sample_weight=any_severe_spinal_weights\n    )\n    condition_losses.append(any_severe_spinal_loss)\n    condition_weights.append(any_severe_scalar)\n    return np.average(condition_losses, weights=condition_weights)\n\n\ndef score_from_train(\n    submission: pd.DataFrame,  # Pass submission.csv as a DataFrame\n    train_df: pd.DataFrame,  # Pass train.csv as a DataFrame\n    row_id_column_name=\"row_id\",\n    any_severe_scalar=1.0,\n    sample_weights: dict[str, int]={\"normal_mild\": 1, \"moderate\": 2, \"severe\": 4},\n):\n    \"\"\"Convert train.csv to sample submission format with Nan values removed\"\"\"\n    target_cols = list(sample_weights.keys())\n    pred = submission.copy() # Copy to prevent changes in original\n    # Normalize values to have a sum of 1.0\n    pred[target_cols] = pred[target_cols].div(pred[target_cols].sum(axis=1), axis=0)\n\n    # Index the study_id in train_df\n    indexed_train_df = train_df.set_index(\"study_id\", verify_integrity=True)\n\n    row_ids = pred[row_id_column_name]\n    study_ids = row_ids.apply(lambda x: x.split('_')[0])\n    locations = row_ids.apply(lambda x: '_'.join(x.split('_')[1:]))\n\n    solution_data = np.zeros_like(pred[target_cols].values)\n    sample_weight_list = []\n    nan_row_ids = set()\n    for idx, (row, study_id, location) in enumerate(zip(row_ids, study_ids, locations)):\n        severity = str(indexed_train_df.at[int(study_id), location]).replace(\"/\", \"_\").lower()\n        if severity in sample_weights:\n            solution_data[idx, target_cols.index(severity)] = 1.0\n            sample_weight_list.append(sample_weights[severity])\n        else:\n            solution_data[idx] = np.nan\n            nan_row_ids.add(row)\n            sample_weight_list.append(np.nan)\n\n    solution = pd.DataFrame({\n        row_id_column_name: pred[row_id_column_name],\n        \"sample_weight\": sample_weight_list\n    })\n    solution[target_cols] = solution_data\n\n    # Change row_ids in nan_row_ids to np.nan\n    pred.loc[pred[row_id_column_name].isin(nan_row_ids), target_cols] = np.nan\n    # Remove nan rows and pass copy to score function\n    # score from https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791 (Version 10)\n    return score(solution.dropna().copy(), pred.dropna().copy(), row_id_column_name, any_severe_scalar)\n```\nCreate a flag for testing on train images: `FAKE_TEST = len(sample_sub) <= 25` this will ensure that during submission of inference notebook the actual test data is used.\n\nIn your notebook swap path to images and other test data dataframes to train path and data. Example:\n```python\nif FAKE_TEST:\n    df = pd.read_csv(DATA_PATH / 'train_series_descriptions.csv')\n    study_ids, base_path = list(df['study_id'].unique()), DATA_PATH / \"train_images\"\n    print(f\"Num of study_ids:\", len(study_ids))\n```\n\nFinally, do the test after `submission.csv` is created:\n```python\nif FAKE_TEST:\n    # Pass submission.csv dataframe and train.csv dataframe\n    print(score_from_train(pd.read_csv(\"submission.csv\"), pd.read_csv(DATA_PATH / \"train.csv\")))\n```\n\nHere are my public notebooks, where I have done this:\n- https://www.kaggle.com/code/coderrkj/lumbar-competition-debugging-notebook\n- https://www.kaggle.com/coderrkj/rsna-pytorch-test-infer\n- https://www.kaggle.com/coderrkj/rsna2024-lsdc-submission-baseline",
      "votes": null
    },
    {
      "id": "2933133",
      "postDate": "07/23/2024 13:50:03",
      "content": "<p>Thank you! I saw your code (before posting the question) and tried to make it work but had some errors. So let's go back to the start:</p>\n<p><em>1_</em> Is my understanding from your comment 'submission and solution' should match (except predicted scores).</p>\n<p><em>2_</em> Challenge public 'test' data has 3 sets: <br>\n<strong>__A</strong> Some 'short' sample <strong>test_series_descriptions.csv</strong><br>\n<strong>__B</strong> Some private test data to calculate Public Scores <strong>(this is what im trying to accomplish)</strong><br>\n<strong>__C</strong> Some final private test data to calculate Private Scores (this will give you the final ranking)</p>\n<p><em>3_</em> Therefore .. how to simulate B and get something as close as posible to B?</p>\n<p>Is your suggestion to predict the whole training set (solution) to get a submission (and then use the score metric)?</p>",
      "rawMarkdown": "Thank you! I saw your code (before posting the question) and tried to make it work but had some errors. So let's go back to the start:\n\n_1__ Is my understanding from your comment 'submission and solution' should match (except predicted scores).\n\n_2__ Challenge public 'test' data has 3 sets: \n____A__ Some 'short' sample **test_series_descriptions.csv**\n____B__ Some private test data to calculate Public Scores **(this is what im trying to accomplish)**\n____C__ Some final private test data to calculate Private Scores (this will give you the final ranking)\n\n_3__ Therefore .. how to simulate B and get something as close as posible to B?\n\nIs your suggestion to predict the whole training set (solution) to get a submission (and then use the score metric)?",
      "votes": null
    },
    {
      "id": "2933157",
      "postDate": "07/23/2024 14:11:35",
      "content": "<p>Is what coderRKJ and me undestood. That you were looking how to validate your validation splits of train.csv with the same metric than the competition. If that's the case, he shared some code about it.</p>\n<p>EDIT: About 2. Challenge has a unique hidden test. We have acces in local to a short example. Your code will execute on full hidden test and will show the score of only a portion of it as public leaderboard. How to obtain insights about hidden is something to I never pay attention. Since there is a reason to beign hidden. Mi suggestion, can be wrong, work with the data provided and may be extra external, but leave hidden as it is.</p>",
      "rawMarkdown": "Is what coderRKJ and me undestood. That you were looking how to validate your validation splits of train.csv with the same metric than the competition. If that's the case, he shared some code about it.\n\nEDIT: About 2. Challenge has a unique hidden test. We have acces in local to a short example. Your code will execute on full hidden test and will show the score of only a portion of it as public leaderboard. How to obtain insights about hidden is something to I never pay attention. Since there is a reason to beign hidden. Mi suggestion, can be wrong, work with the data provided and may be extra external, but leave hidden as it is.",
      "votes": null
    },
    {
      "id": "2933448",
      "postDate": "07/23/2024 17:35:36",
      "content": "<p><a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a>  <a href=\"https://www.kaggle.com/coderrkj\" target=\"_blank\">@coderrkj</a> thanks a lot for your help!</p>\n<p>Based on your code I'm gonna try to score my solution with some 'random training data'.</p>\n<p>I'm running the challenge outside Kaggle cause is faster, but because it takes time to 'upload the code again' Im trying to minimize submissions.</p>\n<p>Thank you for your subjections. Have a great day!</p>",
      "rawMarkdown": "sacuscreed  @coderrkj thanks a lot for your help!\n\nBased on your code I'm gonna try to score my solution with some 'random training data'.\n\nI'm running the challenge outside Kaggle cause is faster, but because it takes time to 'upload the code again' Im trying to minimize submissions.\n\nThank you for your subjections. Have a great day!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2932414,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "07/22/2024 23:06:12",
      "content": "<p>Predict and introduce your \"submission\" (predicted scores) against the \"solution\" (target labels). The weights are 1,2 and 4.<br>\n<code>random_pred = your_prediction</code><br>\nActually random_pred is not random there, is just naive.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2933133,
          "author_name": "karelbecerra",
          "author_url": "",
          "post_date": "07/23/2024 13:50:03",
          "content": "<p>Thank you! I saw your code (before posting the question) and tried to make it work but had some errors. So let's go back to the start:</p>\n<p><em>1_</em> Is my understanding from your comment 'submission and solution' should match (except predicted scores).</p>\n<p><em>2_</em> Challenge public 'test' data has 3 sets: <br>\n<strong>__A</strong> Some 'short' sample <strong>test_series_descriptions.csv</strong><br>\n<strong>__B</strong> Some private test data to calculate Public Scores <strong>(this is what im trying to accomplish)</strong><br>\n<strong>__C</strong> Some final private test data to calculate Private Scores (this will give you the final ranking)</p>\n<p><em>3_</em> Therefore .. how to simulate B and get something as close as posible to B?</p>\n<p>Is your suggestion to predict the whole training set (solution) to get a submission (and then use the score metric)?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2933157,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "07/23/2024 14:11:35",
              "content": "<p>Is what coderRKJ and me undestood. That you were looking how to validate your validation splits of train.csv with the same metric than the competition. If that's the case, he shared some code about it.</p>\n<p>EDIT: About 2. Challenge has a unique hidden test. We have acces in local to a short example. Your code will execute on full hidden test and will show the score of only a portion of it as public leaderboard. How to obtain insights about hidden is something to I never pay attention. Since there is a reason to beign hidden. Mi suggestion, can be wrong, work with the data provided and may be extra external, but leave hidden as it is.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2932540,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "07/23/2024 05:11:17",
      "content": "<p>If your question is how to use the [metric given by the competition host](from <a href=\"https://www.kaggle.com/datasets/seoyunje/kaggle-rsna-2024-metric\" target=\"_blank\">https://www.kaggle.com/datasets/seoyunje/kaggle-rsna-2024-metric</a>) to see the CV score for a test run on train images, then use the following (first part is copied from the metric notebook linked above):</p>\n<pre><code> numpy  np\n pandas  pd\n pandas.api.types\n sklearn.metrics\n\n\n ():\n    \n\n\n () -&gt; :\n    \n     injury_condition  [, , ]:\n         injury_condition  full_location:\n             injury_condition\n     ValueError()\n\n\n () -&gt; :\n    \n\n    target_levels = [, , ]\n\n    \n      pandas.api.types.is_numeric_dtype(submission[target_levels].values):\n         ParticipantVisibleError()\n\n      np.isfinite(submission[target_levels].values).():\n         ParticipantVisibleError()\n\n     solution[target_levels].().() &lt; :\n         ParticipantVisibleError()\n     submission[target_levels].().() &lt; :\n         ParticipantVisibleError()\n\n    solution[] = solution[].apply( x: x.split()[])\n    solution[] = solution[].apply( x: .join(x.split()[:]))\n    solution[] = solution[].apply(get_condition)\n\n     solution[row_id_column_name]\n     submission[row_id_column_name]\n     (submission.columns) == (target_levels)\n\n    submission[] = solution[]\n    submission[] = solution[]\n    submission[] = solution[]\n\n    condition_losses = []\n    condition_weights = []\n     condition  [, , ]:\n        condition_indices = solution.loc[solution[] == condition].index.values\n        condition_loss = sklearn.metrics.log_loss(\n            y_true=solution.loc[condition_indices, target_levels].values,\n            y_pred=submission.loc[condition_indices, target_levels].values,\n            sample_weight=solution.loc[condition_indices, ].values\n        )\n        condition_losses.append(condition_loss)\n        condition_weights.append()\n\n    any_severe_spinal_labels = pd.Series(solution.loc[solution[] == ].groupby()[].())\n    any_severe_spinal_weights = pd.Series(solution.loc[solution[] == ].groupby()[].())\n    any_severe_spinal_predictions = pd.Series(submission.loc[submission[] == ].groupby()[].())\n    any_severe_spinal_loss = sklearn.metrics.log_loss(\n        y_true=any_severe_spinal_labels,\n        y_pred=any_severe_spinal_predictions,\n        sample_weight=any_severe_spinal_weights\n    )\n    condition_losses.append(any_severe_spinal_loss)\n    condition_weights.append(any_severe_scalar)\n     np.average(condition_losses, weights=condition_weights)\n\n\n ():\n    \n    target_cols = (sample_weights.keys())\n    pred = submission.copy() \n    \n    pred[target_cols] = pred[target_cols].div(pred[target_cols].(axis=), axis=)\n\n    \n    indexed_train_df = train_df.set_index(, verify_integrity=)\n\n    row_ids = pred[row_id_column_name]\n    study_ids = row_ids.apply( x: x.split()[])\n    locations = row_ids.apply( x: .join(x.split()[:]))\n\n    solution_data = np.zeros_like(pred[target_cols].values)\n    sample_weight_list = []\n    nan_row_ids = ()\n     idx, (row, study_id, location)  ((row_ids, study_ids, locations)):\n        severity = (indexed_train_df.at[(study_id), location]).replace(, ).lower()\n         severity  sample_weights:\n            solution_data[idx, target_cols.index(severity)] = \n            sample_weight_list.append(sample_weights[severity])\n        :\n            solution_data[idx] = np.nan\n            nan_row_ids.add(row)\n            sample_weight_list.append(np.nan)\n\n    solution = pd.DataFrame({\n        row_id_column_name: pred[row_id_column_name],\n        : sample_weight_list\n    })\n    solution[target_cols] = solution_data\n\n    \n    pred.loc[pred[row_id_column_name].isin(nan_row_ids), target_cols] = np.nan\n    \n    \n     score(solution.dropna().copy(), pred.dropna().copy(), row_id_column_name, any_severe_scalar)\n</code></pre>\n<p>Create a flag for testing on train images: <code>FAKE_TEST = len(sample_sub) &lt;= 25</code> this will ensure that during submission of inference notebook the actual test data is used.</p>\n<p>In your notebook swap path to images and other test data dataframes to train path and data. Example:</p>\n<pre><code> FAKE_TEST:\n    df = pd.read_csv(DATA_PATH / )\n    study_ids, base_path = (df[].unique()), DATA_PATH / \n    (, (study_ids))\n</code></pre>\n<p>Finally, do the test after <code>submission.csv</code> is created:</p>\n<pre><code> FAKE_TEST:\n    \n    (score_from_train(pd.read_csv(), pd.read_csv(DATA_PATH / )))\n</code></pre>\n<p>Here are my public notebooks, where I have done this:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/coderrkj/lumbar-competition-debugging-notebook\" target=\"_blank\">https://www.kaggle.com/code/coderrkj/lumbar-competition-debugging-notebook</a></li>\n<li><a href=\"https://www.kaggle.com/coderrkj/rsna-pytorch-test-infer\" target=\"_blank\">https://www.kaggle.com/coderrkj/rsna-pytorch-test-infer</a></li>\n<li><a href=\"https://www.kaggle.com/coderrkj/rsna2024-lsdc-submission-baseline\" target=\"_blank\">https://www.kaggle.com/coderrkj/rsna2024-lsdc-submission-baseline</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2933448,
      "author_name": "karelbecerra",
      "author_url": "",
      "post_date": "07/23/2024 17:35:36",
      "content": "<p><a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a>  <a href=\"https://www.kaggle.com/coderrkj\" target=\"_blank\">@coderrkj</a> thanks a lot for your help!</p>\n<p>Based on your code I'm gonna try to score my solution with some 'random training data'.</p>\n<p>I'm running the challenge outside Kaggle cause is faster, but because it takes time to 'upload the code again' Im trying to minimize submissions.</p>\n<p>Thank you for your subjections. Have a great day!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2932386": "I did run my notebook (and I can see my public score) so everything works fine. \n\nHowever I'm running my notebook in Visual Studio (locally) and I would like to calculate my public score before submission and double check if there is a significan improvement.\n\nI found this https://www.kaggle.com/code/sadidul012/rsna24-lsdc-validation-scoring-from-trainset as something close to what I want (except submission is random and not predicted). Any clue?",
    "2932414": "Predict and introduce your \"submission\" (predicted scores) against the \"solution\" (target labels). The weights are 1,2 and 4.\n`random_pred = your_prediction`\nActually random_pred is not random there, is just naive.",
    "2932540": "If your question is how to use the [metric given by the competition host](from https://www.kaggle.com/datasets/seoyunje/kaggle-rsna-2024-metric) to see the CV score for a test run on train images, then use the following (first part is copied from the metric notebook linked above):\n\n```python\nimport numpy as np\nimport pandas as pd\nimport pandas.api.types\nimport sklearn.metrics\n\n\nclass ParticipantVisibleError(Exception):\n    pass\n\n\ndef get_condition(full_location: str) -> str:\n    # Given an input like spinal_canal_stenosis_l1_l2 extracts 'spinal'\n    for injury_condition in ['spinal', 'foraminal', 'subarticular']:\n        if injury_condition in full_location:\n            return injury_condition\n    raise ValueError(f'condition not found in {full_location}')\n\n\ndef score(\n        solution: pd.DataFrame,\n        submission: pd.DataFrame,\n        row_id_column_name: str,\n        any_severe_scalar: float\n    ) -> float:\n    '''\n    Pseudocode:\n    1. Calculate the sample weighted log loss for each medical condition:\n    2. Derive a new any_severe label.\n    3. Calculate the sample weighted log loss for the new any_severe label.\n    4. Return the average of all of the label group log losses as the final score, normalized for the number of columns in each group.\n       This mitigates the impact of spinal stenosis having only half as many columns as the other two conditions.\n    '''\n\n    target_levels = ['normal_mild', 'moderate', 'severe']\n\n    # Run basic QC checks on the inputs\n    if not pandas.api.types.is_numeric_dtype(submission[target_levels].values):\n        raise ParticipantVisibleError('All submission values must be numeric')\n\n    if not np.isfinite(submission[target_levels].values).all():\n        raise ParticipantVisibleError('All submission values must be finite')\n\n    if solution[target_levels].min().min() < 0:\n        raise ParticipantVisibleError('All labels must be at least zero')\n    if submission[target_levels].min().min() < 0:\n        raise ParticipantVisibleError('All predictions must be at least zero')\n\n    solution['study_id'] = solution['row_id'].apply(lambda x: x.split('_')[0])\n    solution['location'] = solution['row_id'].apply(lambda x: '_'.join(x.split('_')[1:]))\n    solution['condition'] = solution['row_id'].apply(get_condition)\n\n    del solution[row_id_column_name]\n    del submission[row_id_column_name]\n    assert sorted(submission.columns) == sorted(target_levels)\n\n    submission['study_id'] = solution['study_id']\n    submission['location'] = solution['location']\n    submission['condition'] = solution['condition']\n\n    condition_losses = []\n    condition_weights = []\n    for condition in ['spinal', 'foraminal', 'subarticular']:\n        condition_indices = solution.loc[solution['condition'] == condition].index.values\n        condition_loss = sklearn.metrics.log_loss(\n            y_true=solution.loc[condition_indices, target_levels].values,\n            y_pred=submission.loc[condition_indices, target_levels].values,\n            sample_weight=solution.loc[condition_indices, 'sample_weight'].values\n        )\n        condition_losses.append(condition_loss)\n        condition_weights.append(1)\n\n    any_severe_spinal_labels = pd.Series(solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['severe'].max())\n    any_severe_spinal_weights = pd.Series(solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['sample_weight'].max())\n    any_severe_spinal_predictions = pd.Series(submission.loc[submission['condition'] == 'spinal'].groupby('study_id')['severe'].max())\n    any_severe_spinal_loss = sklearn.metrics.log_loss(\n        y_true=any_severe_spinal_labels,\n        y_pred=any_severe_spinal_predictions,\n        sample_weight=any_severe_spinal_weights\n    )\n    condition_losses.append(any_severe_spinal_loss)\n    condition_weights.append(any_severe_scalar)\n    return np.average(condition_losses, weights=condition_weights)\n\n\ndef score_from_train(\n    submission: pd.DataFrame,  # Pass submission.csv as a DataFrame\n    train_df: pd.DataFrame,  # Pass train.csv as a DataFrame\n    row_id_column_name=\"row_id\",\n    any_severe_scalar=1.0,\n    sample_weights: dict[str, int]={\"normal_mild\": 1, \"moderate\": 2, \"severe\": 4},\n):\n    \"\"\"Convert train.csv to sample submission format with Nan values removed\"\"\"\n    target_cols = list(sample_weights.keys())\n    pred = submission.copy() # Copy to prevent changes in original\n    # Normalize values to have a sum of 1.0\n    pred[target_cols] = pred[target_cols].div(pred[target_cols].sum(axis=1), axis=0)\n\n    # Index the study_id in train_df\n    indexed_train_df = train_df.set_index(\"study_id\", verify_integrity=True)\n\n    row_ids = pred[row_id_column_name]\n    study_ids = row_ids.apply(lambda x: x.split('_')[0])\n    locations = row_ids.apply(lambda x: '_'.join(x.split('_')[1:]))\n\n    solution_data = np.zeros_like(pred[target_cols].values)\n    sample_weight_list = []\n    nan_row_ids = set()\n    for idx, (row, study_id, location) in enumerate(zip(row_ids, study_ids, locations)):\n        severity = str(indexed_train_df.at[int(study_id), location]).replace(\"/\", \"_\").lower()\n        if severity in sample_weights:\n            solution_data[idx, target_cols.index(severity)] = 1.0\n            sample_weight_list.append(sample_weights[severity])\n        else:\n            solution_data[idx] = np.nan\n            nan_row_ids.add(row)\n            sample_weight_list.append(np.nan)\n\n    solution = pd.DataFrame({\n        row_id_column_name: pred[row_id_column_name],\n        \"sample_weight\": sample_weight_list\n    })\n    solution[target_cols] = solution_data\n\n    # Change row_ids in nan_row_ids to np.nan\n    pred.loc[pred[row_id_column_name].isin(nan_row_ids), target_cols] = np.nan\n    # Remove nan rows and pass copy to score function\n    # score from https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549?scriptVersionId=181722791 (Version 10)\n    return score(solution.dropna().copy(), pred.dropna().copy(), row_id_column_name, any_severe_scalar)\n```\nCreate a flag for testing on train images: `FAKE_TEST = len(sample_sub) <= 25` this will ensure that during submission of inference notebook the actual test data is used.\n\nIn your notebook swap path to images and other test data dataframes to train path and data. Example:\n```python\nif FAKE_TEST:\n    df = pd.read_csv(DATA_PATH / 'train_series_descriptions.csv')\n    study_ids, base_path = list(df['study_id'].unique()), DATA_PATH / \"train_images\"\n    print(f\"Num of study_ids:\", len(study_ids))\n```\n\nFinally, do the test after `submission.csv` is created:\n```python\nif FAKE_TEST:\n    # Pass submission.csv dataframe and train.csv dataframe\n    print(score_from_train(pd.read_csv(\"submission.csv\"), pd.read_csv(DATA_PATH / \"train.csv\")))\n```\n\nHere are my public notebooks, where I have done this:\n- https://www.kaggle.com/code/coderrkj/lumbar-competition-debugging-notebook\n- https://www.kaggle.com/coderrkj/rsna-pytorch-test-infer\n- https://www.kaggle.com/coderrkj/rsna2024-lsdc-submission-baseline",
    "2933133": "Thank you! I saw your code (before posting the question) and tried to make it work but had some errors. So let's go back to the start:\n\n_1__ Is my understanding from your comment 'submission and solution' should match (except predicted scores).\n\n_2__ Challenge public 'test' data has 3 sets: \n____A__ Some 'short' sample **test_series_descriptions.csv**\n____B__ Some private test data to calculate Public Scores **(this is what im trying to accomplish)**\n____C__ Some final private test data to calculate Private Scores (this will give you the final ranking)\n\n_3__ Therefore .. how to simulate B and get something as close as posible to B?\n\nIs your suggestion to predict the whole training set (solution) to get a submission (and then use the score metric)?",
    "2933157": "Is what coderRKJ and me undestood. That you were looking how to validate your validation splits of train.csv with the same metric than the competition. If that's the case, he shared some code about it.\n\nEDIT: About 2. Challenge has a unique hidden test. We have acces in local to a short example. Your code will execute on full hidden test and will show the score of only a portion of it as public leaderboard. How to obtain insights about hidden is something to I never pay attention. Since there is a reason to beign hidden. Mi suggestion, can be wrong, work with the data provided and may be extra external, but leave hidden as it is.",
    "2933448": "sacuscreed  @coderrkj thanks a lot for your help!\n\nBased on your code I'm gonna try to score my solution with some 'random training data'.\n\nI'm running the challenge outside Kaggle cause is faster, but because it takes time to 'upload the code again' Im trying to minimize submissions.\n\nThank you for your subjections. Have a great day!"
  },
  "source": "meta"
}