{
  "id": 516848,
  "title": "Does test dataset missing any T1 or T2 series?",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516848",
  "author_name": "",
  "post_date": "2024-07-03T20:09:57.990050200Z",
  "votes": 13,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi <a href=\"https://www.kaggle.com/sohie\" target=\"_blank\">@sohie</a>,<br>\nCan you, please, help to verify if we are missing any sagittal T1 or T2 series for any <code>study_id</code> in the test data (as it happens in training data)?<br>\nI am trying to optimize my inference code. Technically, it can be probed, but the kaggle team feedback is much appreciated.</p>",
  "messages": [
    {
      "id": "2903499",
      "postDate": "07/03/2024 20:09:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohie\" target=\"_blank\">@sohie</a>,<br>\nCan you, please, help to verify if we are missing any sagittal T1 or T2 series for any <code>study_id</code> in the test data (as it happens in training data)?<br>\nI am trying to optimize my inference code. Technically, it can be probed, but the kaggle team feedback is much appreciated.</p>",
      "rawMarkdown": "Hi @sohie,\n\nCan you, please, help to verify if we are missing any sagittal T1 or T2 series for any `study_id` in the test data (as it happens in training data)?\nI am trying to optimize my inference code. Technically, it can be probed, but the kaggle team feedback is much appreciated.",
      "votes": null
    },
    {
      "id": "2906290",
      "postDate": "07/05/2024 13:37:42",
      "content": "<p>I know this does not answer your question, but figured I would share this code in case others find it useful. It basically fills in a dummy prediction for each <code>study_id</code> that you do not predict.</p>\n<pre><code>  import deepcopy\n\ndef insert_missing_rows(sub: pd.DataFrame):\n    if os.getenv():\n        df_all pd.read_csv(\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/test_series_descriptions.csv\")\n    :\n         sub\n\n    label_cols [\n                , , , \n                , , , \n                , , , \n                , , , \n                , , , \n                , , , \n                , , , \n                , , , \n                ,\n            ]\n\n    #  dummy \n     deepcopy(sub.iloc[])\n    row.normal_mild \n    row.moderate \n    row.severe \n\n    #    \n    z sub[\"row_id\"].\n    arr []\n     val  df_all[\"study_id\"].():\n         col  label_cols:\n             \"{}_{}\".format(val, col)\n            if    z:\n                 deepcopy()\n                row.row_id \n                arr.append()\n\n    #   \n    sub pd.concat([sub, pd.DataFrame(arr)], axis, ignore_index)\n    sub sub.sort_values(\"row_id\").reset_index()\n     sub\n\nsub insert_missing_rows(sub)\nsub.to_csv(, index)\n</code></pre>",
      "rawMarkdown": "I know this does not answer your question, but figured I would share this code in case others find it useful. It basically fills in a dummy prediction for each `study_id` that you do not predict.\n\n```\nfrom copy import deepcopy\n\ndef insert_missing_rows(sub: pd.DataFrame):\n    if os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n        df_all= pd.read_csv(\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/test_series_descriptions.csv\")\n    else:\n        return sub\n\n    label_cols= [\n                'spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3', 'spinal_canal_stenosis_l3_l4', \n                'spinal_canal_stenosis_l4_l5', 'spinal_canal_stenosis_l5_s1', 'left_neural_foraminal_narrowing_l1_l2', \n                'left_neural_foraminal_narrowing_l2_l3', 'left_neural_foraminal_narrowing_l3_l4', 'left_neural_foraminal_narrowing_l4_l5', \n                'left_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l1_l2', 'right_neural_foraminal_narrowing_l2_l3', \n                'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l5_s1', \n                'left_subarticular_stenosis_l1_l2', 'left_subarticular_stenosis_l2_l3', 'left_subarticular_stenosis_l3_l4', \n                'left_subarticular_stenosis_l4_l5', 'left_subarticular_stenosis_l5_s1', 'right_subarticular_stenosis_l1_l2', \n                'right_subarticular_stenosis_l2_l3', 'right_subarticular_stenosis_l3_l4', 'right_subarticular_stenosis_l4_l5', \n                'right_subarticular_stenosis_l5_s1',\n            ]\n    \n    # Create dummy row\n    row= deepcopy(sub.iloc[0])\n    row.normal_mild= 1/3\n    row.moderate= 1/3\n    row.severe= 1/3\n    \n    # Check every row exists\n    z= sub[\"row_id\"].values\n    arr= []\n    for val in df_all[\"study_id\"].unique():\n        for col in label_cols:\n            value= \"{}_{}\".format(val, col)\n            if value not in z:\n                row= deepcopy(row)\n                row.row_id= value\n                arr.append(row)\n    \n    # Add new rows\n    sub= pd.concat([sub, pd.DataFrame(arr)], axis=0, ignore_index=True)\n    sub= sub.sort_values(\"row_id\").reset_index(drop=True)\n    return sub\n            \nsub= insert_missing_rows(sub)\nsub.to_csv('submission.csv', index=False)\n```",
      "votes": null
    },
    {
      "id": "2906514",
      "postDate": "07/05/2024 16:24:02",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> thank you, have you tried to probe it with different allocation not (0.33, 0.33, 0.33)?</p>",
      "rawMarkdown": "brendanartley thank you, have you tried to probe it with different allocation not (0.33, 0.33, 0.33)?",
      "votes": null
    },
    {
      "id": "2906624",
      "postDate": "07/05/2024 17:26:07",
      "content": "<p>Nope, we have only tried (0.33, 0.33, 0.33) for now.</p>",
      "rawMarkdown": "Nope, we have only tried (0.33, 0.33, 0.33) for now.",
      "votes": null
    },
    {
      "id": "2906927",
      "postDate": "07/05/2024 21:08:02",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>, why wouldn't you make a prediction on a <code>study_id</code>?</p>",
      "rawMarkdown": "Hello @brendanartley, why wouldn't you make a prediction on a `study_id`?",
      "votes": null
    },
    {
      "id": "2906949",
      "postDate": "07/05/2024 21:40:49",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>. In the training data some study_ids do not have all 3 series types <code>[Sagittal T1, Sagittal T2/STIR, Axial T2]</code>. The study_ids are <code>[2780132468, 3008676218, 2492114990]</code>.</p>\n<p>Let's say I have a pipeline that only makes predictions on <code>Sagittal T1</code> images. In this case, I will not make any predictions on<code>[2780132468, 2492114990]</code>. This code snippet just ensures to fill in those missing study_id's so that the <code>submission.csv</code> is correctly formatted for submission. 🙂</p>",
      "rawMarkdown": "Hi @claverru. In the training data some study_ids do not have all 3 series types `[Sagittal T1, Sagittal T2/STIR, Axial T2]`. The study_ids are `[2780132468, 3008676218, 2492114990]`.\n\nLet's say I have a pipeline that only makes predictions on `Sagittal T1` images. In this case, I will not make any predictions on`[2780132468, 2492114990]`. This code snippet just ensures to fill in those missing study_id's so that the `submission.csv` is correctly formatted for submission. 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2906290,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "07/05/2024 13:37:42",
      "content": "<p>I know this does not answer your question, but figured I would share this code in case others find it useful. It basically fills in a dummy prediction for each <code>study_id</code> that you do not predict.</p>\n<pre><code>  import deepcopy\n\ndef insert_missing_rows(sub: pd.DataFrame):\n    if os.getenv():\n        df_all pd.read_csv(\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/test_series_descriptions.csv\")\n    :\n         sub\n\n    label_cols [\n                , , , \n                , , , \n                , , , \n                , , , \n                , , , \n                , , , \n                , , , \n                , , , \n                ,\n            ]\n\n    #  dummy \n     deepcopy(sub.iloc[])\n    row.normal_mild \n    row.moderate \n    row.severe \n\n    #    \n    z sub[\"row_id\"].\n    arr []\n     val  df_all[\"study_id\"].():\n         col  label_cols:\n             \"{}_{}\".format(val, col)\n            if    z:\n                 deepcopy()\n                row.row_id \n                arr.append()\n\n    #   \n    sub pd.concat([sub, pd.DataFrame(arr)], axis, ignore_index)\n    sub sub.sort_values(\"row_id\").reset_index()\n     sub\n\nsub insert_missing_rows(sub)\nsub.to_csv(, index)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2906514,
          "author_name": "sergiosaharovskiy",
          "author_url": "",
          "post_date": "07/05/2024 16:24:02",
          "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> thank you, have you tried to probe it with different allocation not (0.33, 0.33, 0.33)?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2906624,
              "author_name": "brendanartley",
              "author_url": "",
              "post_date": "07/05/2024 17:26:07",
              "content": "<p>Nope, we have only tried (0.33, 0.33, 0.33) for now.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2906927,
                  "author_name": "claverru",
                  "author_url": "",
                  "post_date": "07/05/2024 21:08:02",
                  "content": "<p>Hello <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>, why wouldn't you make a prediction on a <code>study_id</code>?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2906949,
                      "author_name": "brendanartley",
                      "author_url": "",
                      "post_date": "07/05/2024 21:40:49",
                      "content": "<p>Hi <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>. In the training data some study_ids do not have all 3 series types <code>[Sagittal T1, Sagittal T2/STIR, Axial T2]</code>. The study_ids are <code>[2780132468, 3008676218, 2492114990]</code>.</p>\n<p>Let's say I have a pipeline that only makes predictions on <code>Sagittal T1</code> images. In this case, I will not make any predictions on<code>[2780132468, 2492114990]</code>. This code snippet just ensures to fill in those missing study_id's so that the <code>submission.csv</code> is correctly formatted for submission. 🙂</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2903499": "Hi @sohie,\n\nCan you, please, help to verify if we are missing any sagittal T1 or T2 series for any `study_id` in the test data (as it happens in training data)?\nI am trying to optimize my inference code. Technically, it can be probed, but the kaggle team feedback is much appreciated.",
    "2906290": "I know this does not answer your question, but figured I would share this code in case others find it useful. It basically fills in a dummy prediction for each `study_id` that you do not predict.\n\n```\nfrom copy import deepcopy\n\ndef insert_missing_rows(sub: pd.DataFrame):\n    if os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n        df_all= pd.read_csv(\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/test_series_descriptions.csv\")\n    else:\n        return sub\n\n    label_cols= [\n                'spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3', 'spinal_canal_stenosis_l3_l4', \n                'spinal_canal_stenosis_l4_l5', 'spinal_canal_stenosis_l5_s1', 'left_neural_foraminal_narrowing_l1_l2', \n                'left_neural_foraminal_narrowing_l2_l3', 'left_neural_foraminal_narrowing_l3_l4', 'left_neural_foraminal_narrowing_l4_l5', \n                'left_neural_foraminal_narrowing_l5_s1', 'right_neural_foraminal_narrowing_l1_l2', 'right_neural_foraminal_narrowing_l2_l3', \n                'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l4_l5', 'right_neural_foraminal_narrowing_l5_s1', \n                'left_subarticular_stenosis_l1_l2', 'left_subarticular_stenosis_l2_l3', 'left_subarticular_stenosis_l3_l4', \n                'left_subarticular_stenosis_l4_l5', 'left_subarticular_stenosis_l5_s1', 'right_subarticular_stenosis_l1_l2', \n                'right_subarticular_stenosis_l2_l3', 'right_subarticular_stenosis_l3_l4', 'right_subarticular_stenosis_l4_l5', \n                'right_subarticular_stenosis_l5_s1',\n            ]\n    \n    # Create dummy row\n    row= deepcopy(sub.iloc[0])\n    row.normal_mild= 1/3\n    row.moderate= 1/3\n    row.severe= 1/3\n    \n    # Check every row exists\n    z= sub[\"row_id\"].values\n    arr= []\n    for val in df_all[\"study_id\"].unique():\n        for col in label_cols:\n            value= \"{}_{}\".format(val, col)\n            if value not in z:\n                row= deepcopy(row)\n                row.row_id= value\n                arr.append(row)\n    \n    # Add new rows\n    sub= pd.concat([sub, pd.DataFrame(arr)], axis=0, ignore_index=True)\n    sub= sub.sort_values(\"row_id\").reset_index(drop=True)\n    return sub\n            \nsub= insert_missing_rows(sub)\nsub.to_csv('submission.csv', index=False)\n```",
    "2906514": "brendanartley thank you, have you tried to probe it with different allocation not (0.33, 0.33, 0.33)?",
    "2906624": "Nope, we have only tried (0.33, 0.33, 0.33) for now.",
    "2906927": "Hello @brendanartley, why wouldn't you make a prediction on a `study_id`?",
    "2906949": "Hi @claverru. In the training data some study_ids do not have all 3 series types `[Sagittal T1, Sagittal T2/STIR, Axial T2]`. The study_ids are `[2780132468, 3008676218, 2492114990]`.\n\nLet's say I have a pipeline that only makes predictions on `Sagittal T1` images. In this case, I will not make any predictions on`[2780132468, 2492114990]`. This code snippet just ensures to fill in those missing study_id's so that the `submission.csv` is correctly formatted for submission. 🙂"
  },
  "source": "meta"
}