{
  "id": 515765,
  "title": "Inquiry on Severity Probability Prediction",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/515765",
  "author_name": "",
  "post_date": "2024-06-29T23:52:44.922605100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello everyone,I am wondering about the severity levels: Normal, Mild, Severe, and Moderate. At the end of this computation, we must predict the probability.How can we do that if the training file has 26 features? I am having difficulty understanding this particular part. How can we classify into 3 severity classes and then calculate the probability?</p>",
  "messages": [
    {
      "id": "2896591",
      "postDate": "06/29/2024 23:52:44",
      "content": "<p>Hello everyone,I am wondering about the severity levels: Normal, Mild, Severe, and Moderate. At the end of this computation, we must predict the probability.How can we do that if the training file has 26 features? I am having difficulty understanding this particular part. How can we classify into 3 severity classes and then calculate the probability?</p>",
      "rawMarkdown": "Hello everyone,I am wondering about the severity levels: Normal, Mild, Severe, and Moderate. At the end of this computation, we must predict the probability.How can we do that if the training file has 26 features? I am having difficulty understanding this particular part. How can we classify into 3 severity classes and then calculate the probability?",
      "votes": null
    },
    {
      "id": "2896808",
      "postDate": "06/30/2024 06:20:46",
      "content": "<p>If you take a look at the <code>train.csv</code> file, there are 26 columns, 1 of which is the <code>study_id</code>. So, the remaining columns are: <code>spinal_canal_stenosis_l1_l2</code>, <code>spinal_canal_stenosis_l2_l3</code>, <code>spinal_canal_stenosis_l3_l4</code>, <code>spinal_canal_stenosis_l4_l5</code>, <code>spinal_canal_stenosis_l5_s1</code>, <code>left_neural_foraminal_narrowing_l1_l2</code>, <code>left_neural_foraminal_narrowing_l2_l3</code>, <code>left_neural_foraminal_narrowing_l3_l4</code>, <code>left_neural_foraminal_narrowing_l4_l5</code>, <code>left_neural_foraminal_narrowing_l5_s1</code>, <code>right_neural_foraminal_narrowing_l1_l2</code>, <code>right_neural_foraminal_narrowing_l2_l3</code>, <code>right_neural_foraminal_narrowing_l3_l4</code>, <code>right_neural_foraminal_narrowing_l4_l5</code>, <code>right_neural_foraminal_narrowing_l5_s1</code>, <code>left_subarticular_stenosis_l1_l2</code>, <code>left_subarticular_stenosis_l2_l3</code>, <code>left_subarticular_stenosis_l3_l4</code>, <code>left_subarticular_stenosis_l4_l5</code>, <code>left_subarticular_stenosis_l5_s1</code>, <code>right_subarticular_stenosis_l1_l2</code>, <code>right_subarticular_stenosis_l2_l3</code>, <code>right_subarticular_stenosis_l3_l4</code>, <code>right_subarticular_stenosis_l4_l5</code>, <code>right_subarticular_stenosis_l5_s1</code>.</p>\n<p>If you count the unique values under them, you will get: <code>{'Normal/Mild': 37754, 'Moderate': 7960, 'Severe': 3089, nan: 572}</code></p>\n<p>Ignoring Nan values, for a given study and a location you can build prediction labels. For example, <code>study_id</code>=4003253 and location <code>spinal_canal_stenosis_l1_l2</code> becomes the <code>row_id</code> (in submission.csv) <code>4003253_left_neural_foraminal_narrowing_l1_l2</code></p>\n<pre><code>df_train_main = pd.read_csv(INPUT_DIR / ).set_index()\n(df_train_main.at[, ])\n</code></pre>\n<p>Output: Normal/Mild<br>\nSo, for this <code>row_id</code>, the prediction labels would be:</p>\n<table>\n<thead>\n<tr>\n<th>row_id</th>\n<th>normal_mild</th>\n<th>moderate</th>\n<th>severe</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>4003253_left_neural_foraminal_narrowing_l1_l2</td>\n<td>1.0</td>\n<td>0.0</td>\n<td>0.0</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "If you take a look at the `train.csv` file, there are 26 columns, 1 of which is the `study_id`. So, the remaining columns are: `spinal_canal_stenosis_l1_l2`, `spinal_canal_stenosis_l2_l3`, `spinal_canal_stenosis_l3_l4`, `spinal_canal_stenosis_l4_l5`, `spinal_canal_stenosis_l5_s1`, `left_neural_foraminal_narrowing_l1_l2`, `left_neural_foraminal_narrowing_l2_l3`, `left_neural_foraminal_narrowing_l3_l4`, `left_neural_foraminal_narrowing_l4_l5`, `left_neural_foraminal_narrowing_l5_s1`, `right_neural_foraminal_narrowing_l1_l2`, `right_neural_foraminal_narrowing_l2_l3`, `right_neural_foraminal_narrowing_l3_l4`, `right_neural_foraminal_narrowing_l4_l5`, `right_neural_foraminal_narrowing_l5_s1`, `left_subarticular_stenosis_l1_l2`, `left_subarticular_stenosis_l2_l3`, `left_subarticular_stenosis_l3_l4`, `left_subarticular_stenosis_l4_l5`, `left_subarticular_stenosis_l5_s1`, `right_subarticular_stenosis_l1_l2`, `right_subarticular_stenosis_l2_l3`, `right_subarticular_stenosis_l3_l4`, `right_subarticular_stenosis_l4_l5`, `right_subarticular_stenosis_l5_s1`.\n\nIf you count the unique values under them, you will get: `{'Normal/Mild': 37754, 'Moderate': 7960, 'Severe': 3089, nan: 572}`\n\nIgnoring Nan values, for a given study and a location you can build prediction labels. For example, `study_id`=4003253 and location `spinal_canal_stenosis_l1_l2` becomes the `row_id` (in submission.csv) `4003253_left_neural_foraminal_narrowing_l1_l2`\n\n```python\ndf_train_main = pd.read_csv(INPUT_DIR / 'train.csv').set_index(\"study_id\")\nprint(df_train_main.at[4003253, \"left_neural_foraminal_narrowing_l1_l2\"])\n```\nOutput: Normal/Mild\nSo, for this `row_id`, the prediction labels would be:\n| row_id                                         |   normal_mild |   moderate |   severe |\n|:-----------------------------------------------|--------------:|-----------:|---------:|\n| 4003253_left_neural_foraminal_narrowing_l1_l2 |      1.0 |   0.0 | 0.0 |",
      "votes": null
    },
    {
      "id": "2897225",
      "postDate": "06/30/2024 10:45:05",
      "content": "<p>If I understand correctly, the prediction will be either 1 or 0 for a study_id + 'condition + level'. Is that correct?</p>\n<table>\n<thead>\n<tr>\n<th>row id</th>\n<th>normal_mild</th>\n<th>moderate</th>\n<th>severe</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>4003253_left_neural_foraminal_narrowing_l4_l5</td>\n<td>0</td>\n<td>1</td>\n<td>0</td>\n</tr>\n<tr>\n<td>4646740_spinal_canal_stenosis_l3_l4</td>\n<td>0</td>\n<td>0</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "If I understand correctly, the prediction will be either 1 or 0 for a study_id + 'condition + level'. Is that correct?\n\n|  row id | normal_mild |moderate| severe|\n| --- | --- |\n4003253_left_neural_foraminal_narrowing_l4_l5 | 0 |1|0|\n4646740_spinal_canal_stenosis_l3_l4 | 0 |0|1|",
      "votes": null
    },
    {
      "id": "2897723",
      "postDate": "06/30/2024 17:04:12",
      "content": "<p>Yes, that is correct. These would be the labels.</p>",
      "rawMarkdown": "Yes, that is correct. These would be the labels.",
      "votes": null
    },
    {
      "id": "2897992",
      "postDate": "06/30/2024 20:17:01",
      "content": "<p>Thank you for your guidance</p>",
      "rawMarkdown": "Thank you for your guidance",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2896808,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/30/2024 06:20:46",
      "content": "<p>If you take a look at the <code>train.csv</code> file, there are 26 columns, 1 of which is the <code>study_id</code>. So, the remaining columns are: <code>spinal_canal_stenosis_l1_l2</code>, <code>spinal_canal_stenosis_l2_l3</code>, <code>spinal_canal_stenosis_l3_l4</code>, <code>spinal_canal_stenosis_l4_l5</code>, <code>spinal_canal_stenosis_l5_s1</code>, <code>left_neural_foraminal_narrowing_l1_l2</code>, <code>left_neural_foraminal_narrowing_l2_l3</code>, <code>left_neural_foraminal_narrowing_l3_l4</code>, <code>left_neural_foraminal_narrowing_l4_l5</code>, <code>left_neural_foraminal_narrowing_l5_s1</code>, <code>right_neural_foraminal_narrowing_l1_l2</code>, <code>right_neural_foraminal_narrowing_l2_l3</code>, <code>right_neural_foraminal_narrowing_l3_l4</code>, <code>right_neural_foraminal_narrowing_l4_l5</code>, <code>right_neural_foraminal_narrowing_l5_s1</code>, <code>left_subarticular_stenosis_l1_l2</code>, <code>left_subarticular_stenosis_l2_l3</code>, <code>left_subarticular_stenosis_l3_l4</code>, <code>left_subarticular_stenosis_l4_l5</code>, <code>left_subarticular_stenosis_l5_s1</code>, <code>right_subarticular_stenosis_l1_l2</code>, <code>right_subarticular_stenosis_l2_l3</code>, <code>right_subarticular_stenosis_l3_l4</code>, <code>right_subarticular_stenosis_l4_l5</code>, <code>right_subarticular_stenosis_l5_s1</code>.</p>\n<p>If you count the unique values under them, you will get: <code>{'Normal/Mild': 37754, 'Moderate': 7960, 'Severe': 3089, nan: 572}</code></p>\n<p>Ignoring Nan values, for a given study and a location you can build prediction labels. For example, <code>study_id</code>=4003253 and location <code>spinal_canal_stenosis_l1_l2</code> becomes the <code>row_id</code> (in submission.csv) <code>4003253_left_neural_foraminal_narrowing_l1_l2</code></p>\n<pre><code>df_train_main = pd.read_csv(INPUT_DIR / ).set_index()\n(df_train_main.at[, ])\n</code></pre>\n<p>Output: Normal/Mild<br>\nSo, for this <code>row_id</code>, the prediction labels would be:</p>\n<table>\n<thead>\n<tr>\n<th>row_id</th>\n<th>normal_mild</th>\n<th>moderate</th>\n<th>severe</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>4003253_left_neural_foraminal_narrowing_l1_l2</td>\n<td>1.0</td>\n<td>0.0</td>\n<td>0.0</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 2897225,
          "author_name": "odayahmed",
          "author_url": "",
          "post_date": "06/30/2024 10:45:05",
          "content": "<p>If I understand correctly, the prediction will be either 1 or 0 for a study_id + 'condition + level'. Is that correct?</p>\n<table>\n<thead>\n<tr>\n<th>row id</th>\n<th>normal_mild</th>\n<th>moderate</th>\n<th>severe</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>4003253_left_neural_foraminal_narrowing_l4_l5</td>\n<td>0</td>\n<td>1</td>\n<td>0</td>\n</tr>\n<tr>\n<td>4646740_spinal_canal_stenosis_l3_l4</td>\n<td>0</td>\n<td>0</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>",
          "votes": null,
          "replies": [
            {
              "id": 2897723,
              "author_name": "coderrkj",
              "author_url": "",
              "post_date": "06/30/2024 17:04:12",
              "content": "<p>Yes, that is correct. These would be the labels.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2897992,
                  "author_name": "odayahmed",
                  "author_url": "",
                  "post_date": "06/30/2024 20:17:01",
                  "content": "<p>Thank you for your guidance</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2896591": "Hello everyone,I am wondering about the severity levels: Normal, Mild, Severe, and Moderate. At the end of this computation, we must predict the probability.How can we do that if the training file has 26 features? I am having difficulty understanding this particular part. How can we classify into 3 severity classes and then calculate the probability?",
    "2896808": "If you take a look at the `train.csv` file, there are 26 columns, 1 of which is the `study_id`. So, the remaining columns are: `spinal_canal_stenosis_l1_l2`, `spinal_canal_stenosis_l2_l3`, `spinal_canal_stenosis_l3_l4`, `spinal_canal_stenosis_l4_l5`, `spinal_canal_stenosis_l5_s1`, `left_neural_foraminal_narrowing_l1_l2`, `left_neural_foraminal_narrowing_l2_l3`, `left_neural_foraminal_narrowing_l3_l4`, `left_neural_foraminal_narrowing_l4_l5`, `left_neural_foraminal_narrowing_l5_s1`, `right_neural_foraminal_narrowing_l1_l2`, `right_neural_foraminal_narrowing_l2_l3`, `right_neural_foraminal_narrowing_l3_l4`, `right_neural_foraminal_narrowing_l4_l5`, `right_neural_foraminal_narrowing_l5_s1`, `left_subarticular_stenosis_l1_l2`, `left_subarticular_stenosis_l2_l3`, `left_subarticular_stenosis_l3_l4`, `left_subarticular_stenosis_l4_l5`, `left_subarticular_stenosis_l5_s1`, `right_subarticular_stenosis_l1_l2`, `right_subarticular_stenosis_l2_l3`, `right_subarticular_stenosis_l3_l4`, `right_subarticular_stenosis_l4_l5`, `right_subarticular_stenosis_l5_s1`.\n\nIf you count the unique values under them, you will get: `{'Normal/Mild': 37754, 'Moderate': 7960, 'Severe': 3089, nan: 572}`\n\nIgnoring Nan values, for a given study and a location you can build prediction labels. For example, `study_id`=4003253 and location `spinal_canal_stenosis_l1_l2` becomes the `row_id` (in submission.csv) `4003253_left_neural_foraminal_narrowing_l1_l2`\n\n```python\ndf_train_main = pd.read_csv(INPUT_DIR / 'train.csv').set_index(\"study_id\")\nprint(df_train_main.at[4003253, \"left_neural_foraminal_narrowing_l1_l2\"])\n```\nOutput: Normal/Mild\nSo, for this `row_id`, the prediction labels would be:\n| row_id                                         |   normal_mild |   moderate |   severe |\n|:-----------------------------------------------|--------------:|-----------:|---------:|\n| 4003253_left_neural_foraminal_narrowing_l1_l2 |      1.0 |   0.0 | 0.0 |",
    "2897225": "If I understand correctly, the prediction will be either 1 or 0 for a study_id + 'condition + level'. Is that correct?\n\n|  row id | normal_mild |moderate| severe|\n| --- | --- |\n4003253_left_neural_foraminal_narrowing_l4_l5 | 0 |1|0|\n4646740_spinal_canal_stenosis_l3_l4 | 0 |0|1|",
    "2897723": "Yes, that is correct. These would be the labels.",
    "2897992": "Thank you for your guidance"
  },
  "source": "meta"
}