{
  "id": 512080,
  "title": "Unlabeled conditions in instances",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/512080",
  "author_name": "",
  "post_date": "2024-06-13T11:34:19.799298800Z",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Question to the host <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>I don't know how silly this question is but I haven't found a clear answer anywhere.</p>\n<p>We have many instances in a series, each of them having one or more conditions and levels labeled. </p>\n<p>Can we assume also that, if an instance doesn't have a specific label (e.g. Spinal Canal Stenosis at L1/L2) it means that label is not present in that instance?</p>\n<p>So, every instance that appears in <code>train_label_coordinates.csv</code> is completely labeled (positive and negative)?</p>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": "2870006",
      "postDate": "06/13/2024 11:34:19",
      "content": "<p>Question to the host <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>I don't know how silly this question is but I haven't found a clear answer anywhere.</p>\n<p>We have many instances in a series, each of them having one or more conditions and levels labeled. </p>\n<p>Can we assume also that, if an instance doesn't have a specific label (e.g. Spinal Canal Stenosis at L1/L2) it means that label is not present in that instance?</p>\n<p>So, every instance that appears in <code>train_label_coordinates.csv</code> is completely labeled (positive and negative)?</p>\n<p>Thank you.</p>",
      "rawMarkdown": "Question to the host @sohier \n\nI don't know how silly this question is but I haven't found a clear answer anywhere.\n\nWe have many instances in a series, each of them having one or more conditions and levels labeled. \n\nCan we assume also that, if an instance doesn't have a specific label (e.g. Spinal Canal Stenosis at L1/L2) it means that label is not present in that instance?\n\nSo, every instance that appears in `train_label_coordinates.csv` is completely labeled (positive and negative)?\n\nThank you.",
      "votes": null
    },
    {
      "id": "2870360",
      "postDate": "06/13/2024 14:42:48",
      "content": "<p>From my understanding if <code>NaN</code> is present in <code>train.csv</code>, then that part is missing. I counted if all unique combinations were present (minus columns with <code>NaN</code>) and there where outliers:</p>\n<pre><code>columns = np.asarray(df_train_main.columns)\nset_cols = (columns.tolist()) - {} \n study_id, sub_df  df_train_label.groupby():\n    labeled = (sub_df.condition..lower()..replace(,) +  + sub_df.level..lower()..replace(,)).unique()\n    na_cols = columns[df_train_main[df_train_main.study_id == study_id].isna().values.flat]\n    total = (labeled.tolist() + na_cols.tolist())\n     (total) != :\n        ()\n        ()\n</code></pre>\n<p>Output:</p>\n<pre><code> =    | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n</code></pre>\n<p>So to you question:</p>\n<blockquote>\n  <p>So, every instance that appears in train_label_coordinates.csv is completely labeled (positive and negative)?</p>\n</blockquote>\n<p>These exceptions above exist.</p>",
      "rawMarkdown": "From my understanding if `NaN` is present in `train.csv`, then that part is missing. I counted if all unique combinations were present (minus columns with `NaN`) and there where outliers:\n```python\ncolumns = np.asarray(df_train_main.columns)\nset_cols = set(columns.tolist()) - {\"study_id\"} \nfor study_id, sub_df in df_train_label.groupby(\"study_id\"):\n    labeled = (sub_df.condition.str.lower().str.replace(\" \",\"_\") + \"_\" + sub_df.level.str.lower().str.replace(\"/\",\"_\")).unique()\n    na_cols = columns[df_train_main[df_train_main.study_id == study_id].isna().values.flat]\n    total = set(labeled.tolist() + na_cols.tolist())\n    if len(total) != 25:\n        print(f\"{study_id =:11d} | unique condition and level: {len(labeled)} | missing in train.csv: {len(na_cols)}\")\n        print(f\"\\t{sorted(set_cols - total)}\")\n```\nOutput:\n```test\nstudy_id =   74782131 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  267842058 | unique condition and level: 24 | missing in train.csv: 0\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  267989673 | unique condition and level: 24 | missing in train.csv: 0\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  293713262 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id =  296083289 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  305152236 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  344297746 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  376723024 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id =  390498354 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id =  434488359 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id =  597329259 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  665627263 | unique condition and level: 23 | missing in train.csv: 0\n\t['spinal_canal_stenosis_l4_l5', 'spinal_canal_stenosis_l5_s1']\nstudy_id =  693432872 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  893250212 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id =  934686772 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id =  953218250 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  979209761 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  998688940 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1047914296 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1133158151 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1143209760 | unique condition and level: 23 | missing in train.csv: 1\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1187463765 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1292979992 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1395773918 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1400326269 | unique condition and level: 24 | missing in train.csv: 0\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1431195383 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1452830936 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1557387235 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1567179188 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1613634521 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1681401548 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1722539301 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1745732011 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1868615696 | unique condition and level: 23 | missing in train.csv: 0\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2040217841 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2213304029 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2232794498 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2239199413 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2256339732 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2279142182 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2297295777 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2336516775 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2397650165 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 2548543893 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2566719718 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2615694902 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2839003053 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 2907745008 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2966999234 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3024532039 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l5_s1']\nstudy_id = 3084269121 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3151371929 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3167888497 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3189076268 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3221995449 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3225351618 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l5_s1']\nstudy_id = 3284652867 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3294654272 | unique condition and level: 21 | missing in train.csv: 0\n\t['spinal_canal_stenosis_l2_l3', 'spinal_canal_stenosis_l3_l4', 'spinal_canal_stenosis_l4_l5', 'spinal_canal_stenosis_l5_s1']\nstudy_id = 3303545110 | unique condition and level: 17 | missing in train.csv: 0\n\t['left_subarticular_stenosis_l1_l2', 'left_subarticular_stenosis_l2_l3', 'left_subarticular_stenosis_l3_l4', 'left_subarticular_stenosis_l4_l5', 'right_subarticular_stenosis_l1_l2', 'right_subarticular_stenosis_l2_l3', 'right_subarticular_stenosis_l3_l4', 'right_subarticular_stenosis_l4_l5']\nstudy_id = 3428426893 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3515641631 | unique condition and level: 21 | missing in train.csv: 3\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3525503074 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l2_l3']\nstudy_id = 3537214277 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3674744025 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3711891194 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3824720894 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3850173026 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3906279426 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3936691827 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3942485002 | unique condition and level: 24 | missing in train.csv: 0\n\t['right_neural_foraminal_narrowing_l3_l4']\nstudy_id = 3966998094 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3973705542 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 4072455711 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 4127969449 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 4137194670 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 4146959702 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 4232806580 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\n```\nSo to you question:\n>So, every instance that appears in train_label_coordinates.csv is completely labeled (positive and negative)?\n\nThese exceptions above exist.",
      "votes": null
    },
    {
      "id": "2870408",
      "postDate": "06/13/2024 15:09:35",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/coderrkj\" target=\"_blank\">@coderrkj</a> </p>\n<p>Thanks for the explanation, still, I think it doesn't answer my question. I was asking about condition and level in <code>train_label_coordinates.csv</code>, not about severity in <code>train.csv</code>. I might be missing something though, in that case, let me know. </p>\n<p>For the sake of the question lets say I don't care about severity (Normal/Mild, Moderate, Severe).</p>",
      "rawMarkdown": "Hello @coderrkj \n\nThanks for the explanation, still, I think it doesn't answer my question. I was asking about condition and level in `train_label_coordinates.csv`, not about severity in `train.csv`. I might be missing something though, in that case, let me know. \n\nFor the sake of the question lets say I don't care about severity (Normal/Mild, Moderate, Severe).",
      "votes": null
    },
    {
      "id": "2870474",
      "postDate": "06/13/2024 15:52:49",
      "content": "<p>Sorry, maybe I was not clear. If there is a missing severity in <code>train.csv</code> that means that all the studies CT scan has a missing portion making the condition unmeasurable at that level (that is there should be no point either). </p>\n<p>So, in the above code if that condition x level combination is missing, I ignore it.</p>\n<p>Now getting to my point, even with the above accounted for, there are studies with missing condition x level combination. I have listed the studies and showed which condition and level combination pair is missing.</p>\n<p>E.g. 'spinal_canal_stenosis_l1_l2' =&gt; missing <code>Spinal Canal Stenosis</code> with <code>L1/L2</code> in <code>train_label_coordinates.csv</code></p>",
      "rawMarkdown": "Sorry, maybe I was not clear. If there is a missing severity in `train.csv` that means that all the studies CT scan has a missing portion making the condition unmeasurable at that level (that is there should be no point either). \n\nSo, in the above code if that condition x level combination is missing, I ignore it.\n\nNow getting to my point, even with the above accounted for, there are studies with missing condition x level combination. I have listed the studies and showed which condition and level combination pair is missing.\n\nE.g. 'spinal_canal_stenosis_l1_l2' => missing `Spinal Canal Stenosis` with `L1/L2` in `train_label_coordinates.csv`",
      "votes": null
    },
    {
      "id": "2870862",
      "postDate": "06/13/2024 20:06:21",
      "content": "<p>Take as an example one concrete instance in the dataset:</p>\n<pre><code>                                                                condition_lvl\nstudy_id   series_id  instance_number                                        \n                    Neural Foraminal Narrowing_L1/L2\n                                        Neural Foraminal Narrowing_L2/L3\n                                        Neural Foraminal Narrowing_L3/L4\n                                        Neural Foraminal Narrowing_L4/L5\n                                        Neural Foraminal Narrowing_L5/S1\n                                       Neural Foraminal Narrowing_L1/L2\n                                       Neural Foraminal Narrowing_L2/L3\n                                       Neural Foraminal Narrowing_L3/L4\n                                       Neural Foraminal Narrowing_L4/L5\n                                       Neural Foraminal Narrowing_L5/S1\n</code></pre>\n<p>Is it possible then that it also have a different condition at a different level than the ones that appear in the csv? But for some reason they didn't label it. </p>",
      "rawMarkdown": "Take as an example one concrete instance in the dataset:\n\n```\n                                                                condition_lvl\nstudy_id   series_id  instance_number                                        \n1395773918 2278678071 3                 Left Neural Foraminal Narrowing_L1/L2\n                      3                 Left Neural Foraminal Narrowing_L2/L3\n                      3                 Left Neural Foraminal Narrowing_L3/L4\n                      3                 Left Neural Foraminal Narrowing_L4/L5\n                      3                 Left Neural Foraminal Narrowing_L5/S1\n                      3                Right Neural Foraminal Narrowing_L1/L2\n                      3                Right Neural Foraminal Narrowing_L2/L3\n                      3                Right Neural Foraminal Narrowing_L3/L4\n                      3                Right Neural Foraminal Narrowing_L4/L5\n                      3                Right Neural Foraminal Narrowing_L5/S1\n```\n\nIs it possible then that it also have a different condition at a different level than the ones that appear in the csv? But for some reason they didn't label it.",
      "votes": null
    },
    {
      "id": "2871948",
      "postDate": "06/14/2024 13:32:57",
      "content": "<p>My opinion may be wrong, but after roughly scanning the training dataset, I think that if a patient has some illness, it would be showed in the train.csv and has a level label(mild, moderate or severe). Most of ilness would have corresponding coordinate labels in one or several instance images which could best show the illness. But, I cannot say that <strong>every level label all has coordinate labels</strong>(because mild and normal share the same label) and each ilness has <strong>all the coordinate labels</strong> get correctly annotated. (Maybe just part of them)</p>",
      "rawMarkdown": "My opinion may be wrong, but after roughly scanning the training dataset, I think that if a patient has some illness, it would be showed in the train.csv and has a level label(mild, moderate or severe). Most of ilness would have corresponding coordinate labels in one or several instance images which could best show the illness. But, I cannot say that **every level label all has coordinate labels**(because mild and normal share the same label) and each ilness has **all the coordinate labels** get correctly annotated. (Maybe just part of them)",
      "votes": null
    },
    {
      "id": "2873030",
      "postDate": "06/15/2024 09:15:46",
      "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> I got what you were asking, if there are 2 DICOM instances with different orientations and these 2 slices pass through the same clinical point (say <code>Left Neural Foraminal Narrowing_L1/L2</code>), you are asking if <strong>one will be labelled and the other will not be labelled or if both will be labelled</strong>, right?</p>\n<p>I cannot check this yet as I can't find the relationship between 2 volumes in 3D space. But I can give the following observation:</p>\n<table>\n<thead>\n<tr>\n<th>condition</th>\n<th>Axial T2</th>\n<th>Sagittal T1</th>\n<th>Sagittal T2/STIR</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Left Subarticular Stenosis</td>\n<td>9608</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Right Subarticular Stenosis</td>\n<td>9612</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Left Neural Foraminal Narrowing</td>\n<td>0</td>\n<td>9860</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Right Neural Foraminal Narrowing</td>\n<td>0</td>\n<td>9859</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Spinal Canal Stenosis</td>\n<td>0</td>\n<td>5</td>\n<td>9748</td>\n</tr>\n</tbody>\n</table>\n<p>Subarticular Stenosis is always labelled in Axial T2, Neural Foraminal Narrowing in Sagittal T1 and Spinal Canal Stenosis in Sagittal T2/STIR (except for study_id=3637444890 where all 5 are labelled in Sagittal T1).</p>\n<p>In <code>train_label_coordinates.csv</code>, grouping by <code>study_id</code>, <code>condition</code> and <code>level</code> results in a unique row.</p>\n<p>So, there is a possibility that even through some studies will have multiple Axial T2, Sagittal T1 and Sagittal T2/STIR series, 2 of them may not intersect at these common points and thus have no need to be labelled twice (can't confirm this though)</p>",
      "rawMarkdown": "claverru I got what you were asking, if there are 2 DICOM instances with different orientations and these 2 slices pass through the same clinical point (say `Left Neural Foraminal Narrowing_L1/L2`), you are asking if **one will be labelled and the other will not be labelled or if both will be labelled**, right?\n\nI cannot check this yet as I can't find the relationship between 2 volumes in 3D space. But I can give the following observation:\n| condition                        |   Axial T2 |   Sagittal T1 |   Sagittal T2/STIR |\n|:---------------------------------|-----------:|--------------:|-------------------:|\n| Left Subarticular Stenosis       |       9608 |             0 |                  0 |\n| Right Subarticular Stenosis      |       9612 |             0 |                  0 |\n| Left Neural Foraminal Narrowing  |          0 |          9860 |                  0 |\n| Right Neural Foraminal Narrowing |          0 |          9859 |                  0 |\n| Spinal Canal Stenosis            |          0 |             5 |               9748 |\n\nSubarticular Stenosis is always labelled in Axial T2, Neural Foraminal Narrowing in Sagittal T1 and Spinal Canal Stenosis in Sagittal T2/STIR (except for study_id=3637444890 where all 5 are labelled in Sagittal T1).\n\nIn `train_label_coordinates.csv`, grouping by `study_id`, `condition` and `level` results in a unique row.\n\nSo, there is a possibility that even through some studies will have multiple Axial T2, Sagittal T1 and Sagittal T2/STIR series, 2 of them may not intersect at these common points and thus have no need to be labelled twice (can't confirm this though)",
      "votes": null
    },
    {
      "id": "2874149",
      "postDate": "06/16/2024 04:47:41",
      "content": "<p>Is it possible then that it also have a different condition at a different level than the ones that appear in the csv? But for some reason they didn't label it</p>\n<p>Yes, it can be possible, but unlikely.</p>",
      "rawMarkdown": "Is it possible then that it also have a different condition at a different level than the ones that appear in the csv? But for some reason they didn't label it\n\nYes, it can be possible, but unlikely.",
      "votes": null
    },
    {
      "id": "2874159",
      "postDate": "06/16/2024 04:55:52",
      "content": "<p>Can we assume also that, if an instance doesn't have a specific label (e.g. Spinal Canal Stenosis at L1/L2) it means that label is not present in that instance?</p>\n<p>If you make proper use of the data available, then you will automatically have the question to this question.<br>\nAre you supplying images to your model just with the conditions for each image or also supplying something else? If yes what?</p>",
      "rawMarkdown": "Can we assume also that, if an instance doesn't have a specific label (e.g. Spinal Canal Stenosis at L1/L2) it means that label is not present in that instance?\n\nIf you make proper use of the data available, then you will automatically have the question to this question.\nAre you supplying images to your model just with the conditions for each image or also supplying something else? If yes what?",
      "votes": null
    },
    {
      "id": "2960788",
      "postDate": "08/16/2024 06:28:21",
      "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Hi, have you got the answer for this?  If an instance does not have the annotations in train_label_coordinates.csv, then does it mean that it was not annotated by experts or there is no finding in it -&gt; can be considered Normal/Mild? </p>",
      "rawMarkdown": "claverru Hi, have you got the answer for this?  If an instance does not have the annotations in train_label_coordinates.csv, then does it mean that it was not annotated by experts or there is no finding in it -> can be considered Normal/Mild?",
      "votes": null
    },
    {
      "id": "2960851",
      "postDate": "08/16/2024 08:08:41",
      "content": "<p>before the competition host answers, the best is to treat these samples as \"unlabelled\" (rather than \"negative\")</p>\n<hr>\n<p>data cleaning is part of kaggle competition.<br>\nnote that there are missing labels … AND also wrong labels.</p>\n<p>my suggestion is:</p>\n<ol>\n<li>start by training some models using a subset of full label (and reliable ones)</li>\n<li>use the model to verify the correctness of labels and fill in missing ones.</li>\n<li>with the new data, train your final model. for the modified label (fixed wrong label and missing ones), you can choose to exclude them in loss backward propagation if you don't trust them, etc</li>\n</ol>\n<hr>\n<p>there is yet another method</p>\n<ul>\n<li>focus on self-supervised retraining.</li>\n<li>then you only need a small set of data to fine-tune your classifier.</li>\n</ul>",
      "rawMarkdown": "before the competition host answers, the best is to treat these samples as \"unlabelled\" (rather than \"negative\")\n\n---\n\ndata cleaning is part of kaggle competition.\nnote that there are missing labels ... AND also wrong labels.\n\nmy suggestion is:\n1. start by training some models using a subset of full label (and reliable ones)\n2. use the model to verify the correctness of labels and fill in missing ones.\n3. with the new data, train your final model. for the modified label (fixed wrong label and missing ones), you can choose to exclude them in loss backward propagation if you don't trust them, etc\n\n---\n\nthere is yet another method\n- focus on self-supervised retraining.\n- then you only need a small set of data to fine-tune your classifier.",
      "votes": null
    },
    {
      "id": "2960909",
      "postDate": "08/16/2024 09:25:14",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> well, I made this question more than 2 months ago, when I was starting to play with the data. </p>\n<p>What I understood over the time is that they just labeled images by convenience. So it could happen that they didn't label certain condition in certain image just because maybe from that angle the severity wasn't completely clear. </p>\n<p>Thanks for reaching anyways. And good luck in the rest of the competition. </p>",
      "rawMarkdown": "hengck23 well, I made this question more than 2 months ago, when I was starting to play with the data. \n\nWhat I understood over the time is that they just labeled images by convenience. So it could happen that they didn't label certain condition in certain image just because maybe from that angle the severity wasn't completely clear. \n\nThanks for reaching anyways. And good luck in the rest of the competition.",
      "votes": null
    },
    {
      "id": "2960911",
      "postDate": "08/16/2024 09:26:33",
      "content": "<p><a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a> I answered my conclusion below to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
      "rawMarkdown": "namgalielei I answered my conclusion below to @hengck23",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2870360,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/13/2024 14:42:48",
      "content": "<p>From my understanding if <code>NaN</code> is present in <code>train.csv</code>, then that part is missing. I counted if all unique combinations were present (minus columns with <code>NaN</code>) and there where outliers:</p>\n<pre><code>columns = np.asarray(df_train_main.columns)\nset_cols = (columns.tolist()) - {} \n study_id, sub_df  df_train_label.groupby():\n    labeled = (sub_df.condition..lower()..replace(,) +  + sub_df.level..lower()..replace(,)).unique()\n    na_cols = columns[df_train_main[df_train_main.study_id == study_id].isna().values.flat]\n    total = (labeled.tolist() + na_cols.tolist())\n     (total) != :\n        ()\n        ()\n</code></pre>\n<p>Output:</p>\n<pre><code> =    | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =   | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n =  | unique condition and level:  | missing in train.csv: \n    \n</code></pre>\n<p>So to you question:</p>\n<blockquote>\n  <p>So, every instance that appears in train_label_coordinates.csv is completely labeled (positive and negative)?</p>\n</blockquote>\n<p>These exceptions above exist.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2870408,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "06/13/2024 15:09:35",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/coderrkj\" target=\"_blank\">@coderrkj</a> </p>\n<p>Thanks for the explanation, still, I think it doesn't answer my question. I was asking about condition and level in <code>train_label_coordinates.csv</code>, not about severity in <code>train.csv</code>. I might be missing something though, in that case, let me know. </p>\n<p>For the sake of the question lets say I don't care about severity (Normal/Mild, Moderate, Severe).</p>",
          "votes": null,
          "replies": [
            {
              "id": 2870474,
              "author_name": "coderrkj",
              "author_url": "",
              "post_date": "06/13/2024 15:52:49",
              "content": "<p>Sorry, maybe I was not clear. If there is a missing severity in <code>train.csv</code> that means that all the studies CT scan has a missing portion making the condition unmeasurable at that level (that is there should be no point either). </p>\n<p>So, in the above code if that condition x level combination is missing, I ignore it.</p>\n<p>Now getting to my point, even with the above accounted for, there are studies with missing condition x level combination. I have listed the studies and showed which condition and level combination pair is missing.</p>\n<p>E.g. 'spinal_canal_stenosis_l1_l2' =&gt; missing <code>Spinal Canal Stenosis</code> with <code>L1/L2</code> in <code>train_label_coordinates.csv</code></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2870862,
                  "author_name": "claverru",
                  "author_url": "",
                  "post_date": "06/13/2024 20:06:21",
                  "content": "<p>Take as an example one concrete instance in the dataset:</p>\n<pre><code>                                                                condition_lvl\nstudy_id   series_id  instance_number                                        \n                    Neural Foraminal Narrowing_L1/L2\n                                        Neural Foraminal Narrowing_L2/L3\n                                        Neural Foraminal Narrowing_L3/L4\n                                        Neural Foraminal Narrowing_L4/L5\n                                        Neural Foraminal Narrowing_L5/S1\n                                       Neural Foraminal Narrowing_L1/L2\n                                       Neural Foraminal Narrowing_L2/L3\n                                       Neural Foraminal Narrowing_L3/L4\n                                       Neural Foraminal Narrowing_L4/L5\n                                       Neural Foraminal Narrowing_L5/S1\n</code></pre>\n<p>Is it possible then that it also have a different condition at a different level than the ones that appear in the csv? But for some reason they didn't label it. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2871948,
                      "author_name": "shipuxu",
                      "author_url": "",
                      "post_date": "06/14/2024 13:32:57",
                      "content": "<p>My opinion may be wrong, but after roughly scanning the training dataset, I think that if a patient has some illness, it would be showed in the train.csv and has a level label(mild, moderate or severe). Most of ilness would have corresponding coordinate labels in one or several instance images which could best show the illness. But, I cannot say that <strong>every level label all has coordinate labels</strong>(because mild and normal share the same label) and each ilness has <strong>all the coordinate labels</strong> get correctly annotated. (Maybe just part of them)</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2873030,
                      "author_name": "coderrkj",
                      "author_url": "",
                      "post_date": "06/15/2024 09:15:46",
                      "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> I got what you were asking, if there are 2 DICOM instances with different orientations and these 2 slices pass through the same clinical point (say <code>Left Neural Foraminal Narrowing_L1/L2</code>), you are asking if <strong>one will be labelled and the other will not be labelled or if both will be labelled</strong>, right?</p>\n<p>I cannot check this yet as I can't find the relationship between 2 volumes in 3D space. But I can give the following observation:</p>\n<table>\n<thead>\n<tr>\n<th>condition</th>\n<th>Axial T2</th>\n<th>Sagittal T1</th>\n<th>Sagittal T2/STIR</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Left Subarticular Stenosis</td>\n<td>9608</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Right Subarticular Stenosis</td>\n<td>9612</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Left Neural Foraminal Narrowing</td>\n<td>0</td>\n<td>9860</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Right Neural Foraminal Narrowing</td>\n<td>0</td>\n<td>9859</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Spinal Canal Stenosis</td>\n<td>0</td>\n<td>5</td>\n<td>9748</td>\n</tr>\n</tbody>\n</table>\n<p>Subarticular Stenosis is always labelled in Axial T2, Neural Foraminal Narrowing in Sagittal T1 and Spinal Canal Stenosis in Sagittal T2/STIR (except for study_id=3637444890 where all 5 are labelled in Sagittal T1).</p>\n<p>In <code>train_label_coordinates.csv</code>, grouping by <code>study_id</code>, <code>condition</code> and <code>level</code> results in a unique row.</p>\n<p>So, there is a possibility that even through some studies will have multiple Axial T2, Sagittal T1 and Sagittal T2/STIR series, 2 of them may not intersect at these common points and thus have no need to be labelled twice (can't confirm this though)</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2874149,
                      "author_name": "devsya",
                      "author_url": "",
                      "post_date": "06/16/2024 04:47:41",
                      "content": "<p>Is it possible then that it also have a different condition at a different level than the ones that appear in the csv? But for some reason they didn't label it</p>\n<p>Yes, it can be possible, but unlikely.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2874159,
      "author_name": "devsya",
      "author_url": "",
      "post_date": "06/16/2024 04:55:52",
      "content": "<p>Can we assume also that, if an instance doesn't have a specific label (e.g. Spinal Canal Stenosis at L1/L2) it means that label is not present in that instance?</p>\n<p>If you make proper use of the data available, then you will automatically have the question to this question.<br>\nAre you supplying images to your model just with the conditions for each image or also supplying something else? If yes what?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2960788,
      "author_name": "namgalielei",
      "author_url": "",
      "post_date": "08/16/2024 06:28:21",
      "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Hi, have you got the answer for this?  If an instance does not have the annotations in train_label_coordinates.csv, then does it mean that it was not annotated by experts or there is no finding in it -&gt; can be considered Normal/Mild? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2960911,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "08/16/2024 09:26:33",
          "content": "<p><a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a> I answered my conclusion below to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2960851,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/16/2024 08:08:41",
      "content": "<p>before the competition host answers, the best is to treat these samples as \"unlabelled\" (rather than \"negative\")</p>\n<hr>\n<p>data cleaning is part of kaggle competition.<br>\nnote that there are missing labels … AND also wrong labels.</p>\n<p>my suggestion is:</p>\n<ol>\n<li>start by training some models using a subset of full label (and reliable ones)</li>\n<li>use the model to verify the correctness of labels and fill in missing ones.</li>\n<li>with the new data, train your final model. for the modified label (fixed wrong label and missing ones), you can choose to exclude them in loss backward propagation if you don't trust them, etc</li>\n</ol>\n<hr>\n<p>there is yet another method</p>\n<ul>\n<li>focus on self-supervised retraining.</li>\n<li>then you only need a small set of data to fine-tune your classifier.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2960909,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "08/16/2024 09:25:14",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> well, I made this question more than 2 months ago, when I was starting to play with the data. </p>\n<p>What I understood over the time is that they just labeled images by convenience. So it could happen that they didn't label certain condition in certain image just because maybe from that angle the severity wasn't completely clear. </p>\n<p>Thanks for reaching anyways. And good luck in the rest of the competition. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2870006": "Question to the host @sohier \n\nI don't know how silly this question is but I haven't found a clear answer anywhere.\n\nWe have many instances in a series, each of them having one or more conditions and levels labeled. \n\nCan we assume also that, if an instance doesn't have a specific label (e.g. Spinal Canal Stenosis at L1/L2) it means that label is not present in that instance?\n\nSo, every instance that appears in `train_label_coordinates.csv` is completely labeled (positive and negative)?\n\nThank you.",
    "2870360": "From my understanding if `NaN` is present in `train.csv`, then that part is missing. I counted if all unique combinations were present (minus columns with `NaN`) and there where outliers:\n```python\ncolumns = np.asarray(df_train_main.columns)\nset_cols = set(columns.tolist()) - {\"study_id\"} \nfor study_id, sub_df in df_train_label.groupby(\"study_id\"):\n    labeled = (sub_df.condition.str.lower().str.replace(\" \",\"_\") + \"_\" + sub_df.level.str.lower().str.replace(\"/\",\"_\")).unique()\n    na_cols = columns[df_train_main[df_train_main.study_id == study_id].isna().values.flat]\n    total = set(labeled.tolist() + na_cols.tolist())\n    if len(total) != 25:\n        print(f\"{study_id =:11d} | unique condition and level: {len(labeled)} | missing in train.csv: {len(na_cols)}\")\n        print(f\"\\t{sorted(set_cols - total)}\")\n```\nOutput:\n```test\nstudy_id =   74782131 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  267842058 | unique condition and level: 24 | missing in train.csv: 0\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  267989673 | unique condition and level: 24 | missing in train.csv: 0\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  293713262 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id =  296083289 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  305152236 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  344297746 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  376723024 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id =  390498354 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id =  434488359 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id =  597329259 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  665627263 | unique condition and level: 23 | missing in train.csv: 0\n\t['spinal_canal_stenosis_l4_l5', 'spinal_canal_stenosis_l5_s1']\nstudy_id =  693432872 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  893250212 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id =  934686772 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id =  953218250 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  979209761 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id =  998688940 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1047914296 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1133158151 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1143209760 | unique condition and level: 23 | missing in train.csv: 1\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1187463765 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1292979992 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1395773918 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1400326269 | unique condition and level: 24 | missing in train.csv: 0\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1431195383 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1452830936 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1557387235 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1567179188 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1613634521 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1681401548 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1722539301 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 1745732011 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 1868615696 | unique condition and level: 23 | missing in train.csv: 0\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2040217841 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2213304029 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2232794498 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2239199413 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2256339732 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2279142182 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2297295777 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2336516775 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2397650165 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 2548543893 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2566719718 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2615694902 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2839003053 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 2907745008 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 2966999234 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3024532039 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l5_s1']\nstudy_id = 3084269121 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3151371929 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3167888497 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3189076268 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3221995449 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3225351618 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l5_s1']\nstudy_id = 3284652867 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3294654272 | unique condition and level: 21 | missing in train.csv: 0\n\t['spinal_canal_stenosis_l2_l3', 'spinal_canal_stenosis_l3_l4', 'spinal_canal_stenosis_l4_l5', 'spinal_canal_stenosis_l5_s1']\nstudy_id = 3303545110 | unique condition and level: 17 | missing in train.csv: 0\n\t['left_subarticular_stenosis_l1_l2', 'left_subarticular_stenosis_l2_l3', 'left_subarticular_stenosis_l3_l4', 'left_subarticular_stenosis_l4_l5', 'right_subarticular_stenosis_l1_l2', 'right_subarticular_stenosis_l2_l3', 'right_subarticular_stenosis_l3_l4', 'right_subarticular_stenosis_l4_l5']\nstudy_id = 3428426893 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3515641631 | unique condition and level: 21 | missing in train.csv: 3\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3525503074 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l2_l3']\nstudy_id = 3537214277 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3674744025 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3711891194 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3824720894 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3850173026 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 3906279426 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3936691827 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3942485002 | unique condition and level: 24 | missing in train.csv: 0\n\t['right_neural_foraminal_narrowing_l3_l4']\nstudy_id = 3966998094 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 3973705542 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 4072455711 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\nstudy_id = 4127969449 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 4137194670 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 4146959702 | unique condition and level: 22 | missing in train.csv: 2\n\t['spinal_canal_stenosis_l1_l2']\nstudy_id = 4232806580 | unique condition and level: 19 | missing in train.csv: 4\n\t['spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3']\n```\nSo to you question:\n>So, every instance that appears in train_label_coordinates.csv is completely labeled (positive and negative)?\n\nThese exceptions above exist.",
    "2870408": "Hello @coderrkj \n\nThanks for the explanation, still, I think it doesn't answer my question. I was asking about condition and level in `train_label_coordinates.csv`, not about severity in `train.csv`. I might be missing something though, in that case, let me know. \n\nFor the sake of the question lets say I don't care about severity (Normal/Mild, Moderate, Severe).",
    "2870474": "Sorry, maybe I was not clear. If there is a missing severity in `train.csv` that means that all the studies CT scan has a missing portion making the condition unmeasurable at that level (that is there should be no point either). \n\nSo, in the above code if that condition x level combination is missing, I ignore it.\n\nNow getting to my point, even with the above accounted for, there are studies with missing condition x level combination. I have listed the studies and showed which condition and level combination pair is missing.\n\nE.g. 'spinal_canal_stenosis_l1_l2' => missing `Spinal Canal Stenosis` with `L1/L2` in `train_label_coordinates.csv`",
    "2870862": "Take as an example one concrete instance in the dataset:\n\n```\n                                                                condition_lvl\nstudy_id   series_id  instance_number                                        \n1395773918 2278678071 3                 Left Neural Foraminal Narrowing_L1/L2\n                      3                 Left Neural Foraminal Narrowing_L2/L3\n                      3                 Left Neural Foraminal Narrowing_L3/L4\n                      3                 Left Neural Foraminal Narrowing_L4/L5\n                      3                 Left Neural Foraminal Narrowing_L5/S1\n                      3                Right Neural Foraminal Narrowing_L1/L2\n                      3                Right Neural Foraminal Narrowing_L2/L3\n                      3                Right Neural Foraminal Narrowing_L3/L4\n                      3                Right Neural Foraminal Narrowing_L4/L5\n                      3                Right Neural Foraminal Narrowing_L5/S1\n```\n\nIs it possible then that it also have a different condition at a different level than the ones that appear in the csv? But for some reason they didn't label it.",
    "2871948": "My opinion may be wrong, but after roughly scanning the training dataset, I think that if a patient has some illness, it would be showed in the train.csv and has a level label(mild, moderate or severe). Most of ilness would have corresponding coordinate labels in one or several instance images which could best show the illness. But, I cannot say that **every level label all has coordinate labels**(because mild and normal share the same label) and each ilness has **all the coordinate labels** get correctly annotated. (Maybe just part of them)",
    "2873030": "claverru I got what you were asking, if there are 2 DICOM instances with different orientations and these 2 slices pass through the same clinical point (say `Left Neural Foraminal Narrowing_L1/L2`), you are asking if **one will be labelled and the other will not be labelled or if both will be labelled**, right?\n\nI cannot check this yet as I can't find the relationship between 2 volumes in 3D space. But I can give the following observation:\n| condition                        |   Axial T2 |   Sagittal T1 |   Sagittal T2/STIR |\n|:---------------------------------|-----------:|--------------:|-------------------:|\n| Left Subarticular Stenosis       |       9608 |             0 |                  0 |\n| Right Subarticular Stenosis      |       9612 |             0 |                  0 |\n| Left Neural Foraminal Narrowing  |          0 |          9860 |                  0 |\n| Right Neural Foraminal Narrowing |          0 |          9859 |                  0 |\n| Spinal Canal Stenosis            |          0 |             5 |               9748 |\n\nSubarticular Stenosis is always labelled in Axial T2, Neural Foraminal Narrowing in Sagittal T1 and Spinal Canal Stenosis in Sagittal T2/STIR (except for study_id=3637444890 where all 5 are labelled in Sagittal T1).\n\nIn `train_label_coordinates.csv`, grouping by `study_id`, `condition` and `level` results in a unique row.\n\nSo, there is a possibility that even through some studies will have multiple Axial T2, Sagittal T1 and Sagittal T2/STIR series, 2 of them may not intersect at these common points and thus have no need to be labelled twice (can't confirm this though)",
    "2874149": "Is it possible then that it also have a different condition at a different level than the ones that appear in the csv? But for some reason they didn't label it\n\nYes, it can be possible, but unlikely.",
    "2874159": "Can we assume also that, if an instance doesn't have a specific label (e.g. Spinal Canal Stenosis at L1/L2) it means that label is not present in that instance?\n\nIf you make proper use of the data available, then you will automatically have the question to this question.\nAre you supplying images to your model just with the conditions for each image or also supplying something else? If yes what?",
    "2960788": "claverru Hi, have you got the answer for this?  If an instance does not have the annotations in train_label_coordinates.csv, then does it mean that it was not annotated by experts or there is no finding in it -> can be considered Normal/Mild?",
    "2960851": "before the competition host answers, the best is to treat these samples as \"unlabelled\" (rather than \"negative\")\n\n---\n\ndata cleaning is part of kaggle competition.\nnote that there are missing labels ... AND also wrong labels.\n\nmy suggestion is:\n1. start by training some models using a subset of full label (and reliable ones)\n2. use the model to verify the correctness of labels and fill in missing ones.\n3. with the new data, train your final model. for the modified label (fixed wrong label and missing ones), you can choose to exclude them in loss backward propagation if you don't trust them, etc\n\n---\n\nthere is yet another method\n- focus on self-supervised retraining.\n- then you only need a small set of data to fine-tune your classifier.",
    "2960909": "hengck23 well, I made this question more than 2 months ago, when I was starting to play with the data. \n\nWhat I understood over the time is that they just labeled images by convenience. So it could happen that they didn't label certain condition in certain image just because maybe from that angle the severity wasn't completely clear. \n\nThanks for reaching anyways. And good luck in the rest of the competition.",
    "2960911": "namgalielei I answered my conclusion below to @hengck23"
  },
  "source": "meta"
}