{
  "id": 513909,
  "title": "Null 'nan' values in train.csv? What do they mean?",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/513909",
  "author_name": "",
  "post_date": "2024-06-22T04:38:20.938350500Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I am currently doing <strong>EDA</strong> over train.csv and I found that there are a lot of missing values in the columns. What do these missing values mean? </p>\n<p>Do they mean that, The labellers were not sure of a given stenosis conditon (<em>out of the 5 available</em>) and a given lumbar region / level (<em>out of the 5 available</em>) <strong>OR</strong> is there any other reason.</p>\n<p>Is there any official notice on these <code>nan</code> values.</p>",
  "messages": [
    {
      "id": "2883637",
      "postDate": "06/22/2024 04:38:20",
      "content": "<p>Hello everyone,</p>\n<p>I am currently doing <strong>EDA</strong> over train.csv and I found that there are a lot of missing values in the columns. What do these missing values mean? </p>\n<p>Do they mean that, The labellers were not sure of a given stenosis conditon (<em>out of the 5 available</em>) and a given lumbar region / level (<em>out of the 5 available</em>) <strong>OR</strong> is there any other reason.</p>\n<p>Is there any official notice on these <code>nan</code> values.</p>",
      "rawMarkdown": "Hello everyone,\n\nI am currently doing **EDA** over train.csv and I found that there are a lot of missing values in the columns. What do these missing values mean? \n\nDo they mean that, The labellers were not sure of a given stenosis conditon (*out of the 5 available*) and a given lumbar region / level (*out of the 5 available*) **OR** is there any other reason.\n\nIs there any official notice on these `nan` values.",
      "votes": null
    },
    {
      "id": "2884047",
      "postDate": "06/22/2024 09:17:55",
      "content": "<p>Initially, I thought that <code>nan</code> values in <code>train.csv</code> were the cases where regions are absent from the MRI images. This was because in the <a href=\"www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/overview/evaluation\" target=\"_blank\">Evaluation</a>, it was written:</p>\n<blockquote>\n  <p>In rare cases the lowest vertebrae aren't visible in the imagery.</p>\n</blockquote>\n<p>But on your question, I looked into it to verify that was the case in the metadata given.</p>\n<pre><code> pathlib  Path\n\n numpy  np\n pandas  pd\n\nINPUT_DIR = Path()\ndf_train_main = pd.read_csv(INPUT_DIR / )\ndf_train_label = pd.read_csv(INPUT_DIR / )\n\ncolumns = np.asarray(df_train_main.columns)\n study_id  indexed_df_train[indexed_df_train.isna().(axis=)].index:\n    sub_df = df_train_label[df_train_label.study_id == study_id]\n    labeled = (sub_df.condition..lower()..replace(,) +  + sub_df.level..lower()..replace(,)).unique()\n    na_cols = columns[df_train_main[df_train_main.study_id == study_id].isna().values.flat]\n    common = (labeled) &amp; (na_cols)\n     common:\n        (study_id, common)\n</code></pre>\n<p>Output:</p>\n<pre><code>462494704 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n482624307 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n779042451 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n1524089207 {'left_subarticular_stenosis_l3_l4', 'left_subarticular_stenosis_l1_l2', 'left_subarticular_stenosis_l4_l5', 'left_subarticular_stenosis_l5_s1', 'left_subarticular_stenosis_l2_l3'}\n2507107985 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n2607462358 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n3781188430 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n</code></pre>\n<p>As you can see above, out of the 185 studies containing <code>nan</code> values, 7 of them contain x, y labelling in <code>train_label_coordinates.csv</code> where the columns were marked as <code>nan</code> in <code>train.csv</code>. This seems to be specific to <code>neural_foraminal_narrowing</code> type of labels.</p>\n<p>This conflicts with \"the volume did not contain the region\" and \"labelers could not locate the level or region\" possibilities for the meaning of <code>nan</code> values. So maybe your \"labelers were not sure of a given condition\" maybe the best possibility.</p>",
      "rawMarkdown": "Initially, I thought that `nan` values in `train.csv` were the cases where regions are absent from the MRI images. This was because in the [Evaluation](www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/overview/evaluation), it was written:\n> In rare cases the lowest vertebrae aren't visible in the imagery.\n\nBut on your question, I looked into it to verify that was the case in the metadata given.\n\n```python\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\n\nINPUT_DIR = Path(\"../input/rsna-2024-lumbar-spine-degenerative-classification\")\ndf_train_main = pd.read_csv(INPUT_DIR / 'train.csv')\ndf_train_label = pd.read_csv(INPUT_DIR / 'train_label_coordinates.csv')\n\ncolumns = np.asarray(df_train_main.columns)\nfor study_id in indexed_df_train[indexed_df_train.isna().any(axis=1)].index:\n    sub_df = df_train_label[df_train_label.study_id == study_id]\n    labeled = (sub_df.condition.str.lower().str.replace(\" \",\"_\") + \"_\" + sub_df.level.str.lower().str.replace(\"/\",\"_\")).unique()\n    na_cols = columns[df_train_main[df_train_main.study_id == study_id].isna().values.flat]\n    common = set(labeled) & set(na_cols)\n    if common:\n        print(study_id, common)\n```\nOutput:\n```text\n462494704 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n482624307 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n779042451 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n1524089207 {'left_subarticular_stenosis_l3_l4', 'left_subarticular_stenosis_l1_l2', 'left_subarticular_stenosis_l4_l5', 'left_subarticular_stenosis_l5_s1', 'left_subarticular_stenosis_l2_l3'}\n2507107985 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n2607462358 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n3781188430 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n```\nAs you can see above, out of the 185 studies containing `nan` values, 7 of them contain x, y labelling in `train_label_coordinates.csv` where the columns were marked as `nan` in `train.csv`. This seems to be specific to `neural_foraminal_narrowing` type of labels.\n\nThis conflicts with \"the volume did not contain the region\" and \"labelers could not locate the level or region\" possibilities for the meaning of `nan` values. So maybe your \"labelers were not sure of a given condition\" maybe the best possibility.",
      "votes": null
    },
    {
      "id": "2891542",
      "postDate": "06/26/2024 17:09:46",
      "content": "<p>They mostly mean it is not in view of the image taken or that there was uncertainty.  There is official notice on the overview of this competition.  &gt;<em>In rare cases the lowest vertebrae aren't visible in the imagery. You still need to make predictions (nulls will cause errors), but those rows will not be scored.</em>.  I recommend handling this by setting the labels for nan values to -100 as CrossEntropyLoss will ignore this and it will not negatively impact training your model.  Additionally, these nulls are not scored in the weighted log loss metric.  I am happy to clarify more if you have more questions, I typed this up very quick and probably made 0 sense.</p>",
      "rawMarkdown": "They mostly mean it is not in view of the image taken or that there was uncertainty.  There is official notice on the overview of this competition.  >*In rare cases the lowest vertebrae aren't visible in the imagery. You still need to make predictions (nulls will cause errors), but those rows will not be scored.*.  I recommend handling this by setting the labels for nan values to -100 as CrossEntropyLoss will ignore this and it will not negatively impact training your model.  Additionally, these nulls are not scored in the weighted log loss metric.  I am happy to clarify more if you have more questions, I typed this up very quick and probably made 0 sense.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2884047,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/22/2024 09:17:55",
      "content": "<p>Initially, I thought that <code>nan</code> values in <code>train.csv</code> were the cases where regions are absent from the MRI images. This was because in the <a href=\"www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/overview/evaluation\" target=\"_blank\">Evaluation</a>, it was written:</p>\n<blockquote>\n  <p>In rare cases the lowest vertebrae aren't visible in the imagery.</p>\n</blockquote>\n<p>But on your question, I looked into it to verify that was the case in the metadata given.</p>\n<pre><code> pathlib  Path\n\n numpy  np\n pandas  pd\n\nINPUT_DIR = Path()\ndf_train_main = pd.read_csv(INPUT_DIR / )\ndf_train_label = pd.read_csv(INPUT_DIR / )\n\ncolumns = np.asarray(df_train_main.columns)\n study_id  indexed_df_train[indexed_df_train.isna().(axis=)].index:\n    sub_df = df_train_label[df_train_label.study_id == study_id]\n    labeled = (sub_df.condition..lower()..replace(,) +  + sub_df.level..lower()..replace(,)).unique()\n    na_cols = columns[df_train_main[df_train_main.study_id == study_id].isna().values.flat]\n    common = (labeled) &amp; (na_cols)\n     common:\n        (study_id, common)\n</code></pre>\n<p>Output:</p>\n<pre><code>462494704 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n482624307 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n779042451 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n1524089207 {'left_subarticular_stenosis_l3_l4', 'left_subarticular_stenosis_l1_l2', 'left_subarticular_stenosis_l4_l5', 'left_subarticular_stenosis_l5_s1', 'left_subarticular_stenosis_l2_l3'}\n2507107985 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n2607462358 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n3781188430 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n</code></pre>\n<p>As you can see above, out of the 185 studies containing <code>nan</code> values, 7 of them contain x, y labelling in <code>train_label_coordinates.csv</code> where the columns were marked as <code>nan</code> in <code>train.csv</code>. This seems to be specific to <code>neural_foraminal_narrowing</code> type of labels.</p>\n<p>This conflicts with \"the volume did not contain the region\" and \"labelers could not locate the level or region\" possibilities for the meaning of <code>nan</code> values. So maybe your \"labelers were not sure of a given condition\" maybe the best possibility.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2891542,
      "author_name": "connorjd",
      "author_url": "",
      "post_date": "06/26/2024 17:09:46",
      "content": "<p>They mostly mean it is not in view of the image taken or that there was uncertainty.  There is official notice on the overview of this competition.  &gt;<em>In rare cases the lowest vertebrae aren't visible in the imagery. You still need to make predictions (nulls will cause errors), but those rows will not be scored.</em>.  I recommend handling this by setting the labels for nan values to -100 as CrossEntropyLoss will ignore this and it will not negatively impact training your model.  Additionally, these nulls are not scored in the weighted log loss metric.  I am happy to clarify more if you have more questions, I typed this up very quick and probably made 0 sense.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2883637": "Hello everyone,\n\nI am currently doing **EDA** over train.csv and I found that there are a lot of missing values in the columns. What do these missing values mean? \n\nDo they mean that, The labellers were not sure of a given stenosis conditon (*out of the 5 available*) and a given lumbar region / level (*out of the 5 available*) **OR** is there any other reason.\n\nIs there any official notice on these `nan` values.",
    "2884047": "Initially, I thought that `nan` values in `train.csv` were the cases where regions are absent from the MRI images. This was because in the [Evaluation](www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/overview/evaluation), it was written:\n> In rare cases the lowest vertebrae aren't visible in the imagery.\n\nBut on your question, I looked into it to verify that was the case in the metadata given.\n\n```python\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\n\nINPUT_DIR = Path(\"../input/rsna-2024-lumbar-spine-degenerative-classification\")\ndf_train_main = pd.read_csv(INPUT_DIR / 'train.csv')\ndf_train_label = pd.read_csv(INPUT_DIR / 'train_label_coordinates.csv')\n\ncolumns = np.asarray(df_train_main.columns)\nfor study_id in indexed_df_train[indexed_df_train.isna().any(axis=1)].index:\n    sub_df = df_train_label[df_train_label.study_id == study_id]\n    labeled = (sub_df.condition.str.lower().str.replace(\" \",\"_\") + \"_\" + sub_df.level.str.lower().str.replace(\"/\",\"_\")).unique()\n    na_cols = columns[df_train_main[df_train_main.study_id == study_id].isna().values.flat]\n    common = set(labeled) & set(na_cols)\n    if common:\n        print(study_id, common)\n```\nOutput:\n```text\n462494704 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n482624307 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n779042451 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n1524089207 {'left_subarticular_stenosis_l3_l4', 'left_subarticular_stenosis_l1_l2', 'left_subarticular_stenosis_l4_l5', 'left_subarticular_stenosis_l5_s1', 'left_subarticular_stenosis_l2_l3'}\n2507107985 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n2607462358 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n3781188430 {'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l2_l3', 'right_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l1_l2'}\n```\nAs you can see above, out of the 185 studies containing `nan` values, 7 of them contain x, y labelling in `train_label_coordinates.csv` where the columns were marked as `nan` in `train.csv`. This seems to be specific to `neural_foraminal_narrowing` type of labels.\n\nThis conflicts with \"the volume did not contain the region\" and \"labelers could not locate the level or region\" possibilities for the meaning of `nan` values. So maybe your \"labelers were not sure of a given condition\" maybe the best possibility.",
    "2891542": "They mostly mean it is not in view of the image taken or that there was uncertainty.  There is official notice on the overview of this competition.  >*In rare cases the lowest vertebrae aren't visible in the imagery. You still need to make predictions (nulls will cause errors), but those rows will not be scored.*.  I recommend handling this by setting the labels for nan values to -100 as CrossEntropyLoss will ignore this and it will not negatively impact training your model.  Additionally, these nulls are not scored in the weighted log loss metric.  I am happy to clarify more if you have more questions, I typed this up very quick and probably made 0 sense."
  },
  "source": "meta"
}