{
  "id": 514891,
  "title": "Some images are of the cervical vertebrae, not of the lumbar spine?",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/514891",
  "author_name": "",
  "post_date": "2024-06-26T03:33:40.729304700Z",
  "votes": 15,
  "comment_count": 11,
  "views": 0,
  "content": "<p>please take a look:    study_id: 3637444890     series_id: 3892989905</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F0f158495a8ffe4f4d2368a52d3406694%2F006.png?generation=1719372819072726&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2890264",
      "postDate": "06/26/2024 03:33:40",
      "content": "<p>please take a look:    study_id: 3637444890     series_id: 3892989905</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F0f158495a8ffe4f4d2368a52d3406694%2F006.png?generation=1719372819072726&amp;alt=media\"></p>",
      "rawMarkdown": "please take a look:    study_id: 3637444890     series_id: 3892989905\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F0f158495a8ffe4f4d2368a52d3406694%2F006.png?generation=1719372819072726&alt=media)",
      "votes": null
    },
    {
      "id": "2890326",
      "postDate": "06/26/2024 04:29:33",
      "content": "<p>Yes, I can confirm that this is a Sagittal slice of the upper body. The cerebellum can be seen at the top and there are no lumbar vertebrae in the bottom part of the image. Thanks for pointing this case out.</p>\n<p>Fun fact: study_id 3637444890 is the only one where Sagittal T1 is used to label Spinal Canal Stenosis. All other use Sagittal T2/STIR. I guess this is the reason.</p>",
      "rawMarkdown": "Yes, I can confirm that this is a Sagittal slice of the upper body. The cerebellum can be seen at the top and there are no lumbar vertebrae in the bottom part of the image. Thanks for pointing this case out.\n\nFun fact: study_id 3637444890 is the only one where Sagittal T1 is used to label Spinal Canal Stenosis. All other use Sagittal T2/STIR. I guess this is the reason.",
      "votes": null
    },
    {
      "id": "2890883",
      "postDate": "06/26/2024 11:09:20",
      "content": "<p>Yes, it is the cervical spine image, I was about to post it few days ago. Just drop it. I have not found any other incosistencies of this nature anywhere else.</p>\n<p>The algorithm your teammate has proposed is actually adds it to the input channels during the dataloading. I have not noticed a big difference in the score though with/or without it.</p>",
      "rawMarkdown": "Yes, it is the cervical spine image, I was about to post it few days ago. Just drop it. I have not found any other incosistencies of this nature anywhere else.\n\nThe algorithm your teammate has proposed is actually adds it to the input channels during the dataloading. I have not noticed a big difference in the score though with/or without it.",
      "votes": null
    },
    {
      "id": "2891500",
      "postDate": "06/26/2024 16:57:10",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a>. How did you identify this?</p>",
      "rawMarkdown": "Thanks for sharing @lihaoweicvch. How did you identify this?",
      "votes": null
    },
    {
      "id": "2891625",
      "postDate": "06/26/2024 18:23:50",
      "content": "<p>Idk how the author  identified that and you did not ask me, though if you let me it might help answer other people questions. If you go by simple merge <code>test_series_descriptions.csv</code> and <code>train_label_coordinates.csv</code> and sort by missing coordinates - there will be some hits. I also went thru all sagittal examples, I did not find anything else like this.</p>\n<pre><code>mrgd = train_desc.merge(train_coord, on=[, ], how=)\n(mrgd[mrgd.x.isna()].to_string())\n</code></pre>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>study_id</th>\n<th>series_id</th>\n<th>series_description</th>\n<th>instance_number</th>\n<th>condition</th>\n<th>level</th>\n<th>x</th>\n<th>y</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>33978</td>\n<td>3008676218</td>\n<td>542282425</td>\n<td>Sagittal T1</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n</tr>\n<tr>\n<td>33979</td>\n<td>3008676218</td>\n<td>3636216534</td>\n<td>Axial T2</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n</tr>\n<tr>\n<td>41256</td>\n<td>3637444890</td>\n<td>3892989905</td>\n<td>Sagittal T2/STIR</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Idk how the author  identified that and you did not ask me, though if you let me it might help answer other people questions. If you go by simple merge `test_series_descriptions.csv` and `train_label_coordinates.csv` and sort by missing coordinates - there will be some hits. I also went thru all sagittal examples, I did not find anything else like this.\n```python\nmrgd = train_desc.merge(train_coord, on=['study_id', 'series_id'], how='left')\nprint(mrgd[mrgd.x.isna()].to_string())\n```\n\n|    | study_id   | series_id   | series_description | instance_number | condition | level | x   | y   |\n|----|------------|-------------|--------------------|-----------------|-----------|-------|-----|-----|\n| 33978 | 3008676218 | 542282425   | Sagittal T1         | NaN             | NaN       | NaN   | NaN | NaN |\n| 33979 | 3008676218 | 3636216534  | Axial T2            | NaN             | NaN       | NaN   | NaN | NaN |\n| 41256 | 3637444890 | 3892989905  | Sagittal T2/STIR    | NaN             | NaN       | NaN   | NaN | NaN |",
      "votes": null
    },
    {
      "id": "2891730",
      "postDate": "06/26/2024 20:02:12",
      "content": "<p>My fault, <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>. I should have asked you both!</p>",
      "rawMarkdown": "My fault, @sergiosaharovskiy. I should have asked you both!",
      "votes": null
    },
    {
      "id": "2896221",
      "postDate": "06/29/2024 17:10:35",
      "content": "<p>Thank you for this valuable insight.  <br>\nI have noticed that total_train_label_cords_desc['instance_number'][41256] produces 5 and total_train_label_cords_desc['condition'][41256] produces 'Right Neural Foraminal Narrowing' and not nan. I am a bit confused by this output has the dataset been modified?</p>\n<p>Your responses will be highly valuable. </p>",
      "rawMarkdown": "Thank you for this valuable insight.  \nI have noticed that total_train_label_cords_desc['instance_number'][41256] produces 5 and total_train_label_cords_desc['condition'][41256] produces 'Right Neural Foraminal Narrowing' and not nan. I am a bit confused by this output has the dataset been modified?\n\nYour responses will be highly valuable.",
      "votes": null
    },
    {
      "id": "2896316",
      "postDate": "06/29/2024 18:41:34",
      "content": "<p><a href=\"https://www.kaggle.com/devsya\" target=\"_blank\">@devsya</a>, it would be great if you could show a snippet on how you got <code>total_train_label_cords_desc</code>. <br>\nI reran the code Sergey gave in a Kaggle notebook cell:</p>\n<pre><code> pathlib  Path\n pandas  pd\n\nINPUT_DIR = Path()\ndf_train_label = pd.read_csv(INPUT_DIR / )\ndf_train_desc = pd.read_csv(INPUT_DIR / )\n\nmrgd = df_train_desc.merge(df_train_label, on=[, ], how=)\nmrgd[mrgd.x.isna()]\n</code></pre>\n<p>and the result is the same as above. To be specific to your question, <code>mrgd.loc[41256]</code> gives:</p>\n<pre><code>study_id                    3637444890\nseries_id                   3892989905\nseries_description    Sagittal T2/STIR\ninstance_number                    NaN\ncondition                          NaN\nlevel                              NaN\nx                                  NaN\ny                                  NaN\nName: 41256, dtype: object\n</code></pre>",
      "rawMarkdown": "devsya, it would be great if you could show a snippet on how you got `total_train_label_cords_desc`. \nI reran the code Sergey gave in a Kaggle notebook cell:\n```python\nfrom pathlib import Path\nimport pandas as pd\n\nINPUT_DIR = Path(\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification\")\ndf_train_label = pd.read_csv(INPUT_DIR / 'train_label_coordinates.csv')\ndf_train_desc = pd.read_csv(INPUT_DIR / 'train_series_descriptions.csv')\n\nmrgd = df_train_desc.merge(df_train_label, on=['study_id', 'series_id'], how='left')\nmrgd[mrgd.x.isna()]\n```\nand the result is the same as above. To be specific to your question, `mrgd.loc[41256]` gives:\n```text\nstudy_id                    3637444890\nseries_id                   3892989905\nseries_description    Sagittal T2/STIR\ninstance_number                    NaN\ncondition                          NaN\nlevel                              NaN\nx                                  NaN\ny                                  NaN\nName: 41256, dtype: object\n```",
      "votes": null
    },
    {
      "id": "2896799",
      "postDate": "06/30/2024 06:12:09",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14118568%2Fae6104f614419dae16aa9efd654664e5%2Fscreenshot_train_label_cords.png?generation=1719727922004953&amp;alt=media\" alt=\"See This\"></p>",
      "rawMarkdown": "![See This](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14118568%2Fae6104f614419dae16aa9efd654664e5%2Fscreenshot_train_label_cords.png?generation=1719727922004953&alt=media)",
      "votes": null
    },
    {
      "id": "2896806",
      "postDate": "06/30/2024 06:20:09",
      "content": "<p><a href=\"https://www.kaggle.com/coderrkj\" target=\"_blank\">@coderrkj</a> In the image you would see that 41256 points to a different series ID. So I tried this : train_label_cords[(train_label_cords['study_id']=='3637444890') &amp; (train_label_cords['series_id']=='3892989905')]<br>\nAnd it returned zero rows. I think that the dataset has been updated. You should also check it by clicking on the three dots of the main folder of the competition dataset and then choosing the check for updates option.</p>",
      "rawMarkdown": "coderrkj In the image you would see that 41256 points to a different series ID. So I tried this : train_label_cords[(train_label_cords['study_id']=='3637444890') & (train_label_cords['series_id']=='3892989905')]\nAnd it returned zero rows. I think that the dataset has been updated. You should also check it by clicking on the three dots of the main folder of the competition dataset and then choosing the check for updates option.",
      "votes": null
    },
    {
      "id": "2896964",
      "postDate": "06/30/2024 07:44:58",
      "content": "<p>The confusion that you are having is that index number <code>41256</code> in <code>df_train_label</code> and <code>df_train_desc.merge(df_train_label, on=['study_id', 'series_id'], how='left')</code> refer to 2 separate things.<br>\nFirst on is just the 41256th row in <code>train_label_coordinates.csv</code> the other is the 41256th row in <strong>left SQL join</strong> of <code>train_series_descriptions.csv</code> and <code>train_label_coordinates.csv</code> on the columns <code>study_id</code> and <code>series_id</code>.<br>\nIf you do <code>print(len(mrgd), len(df_train_label))</code> from my code cell that I gave above, you get <code>48695 48692</code>. There are 3 series missing in <code>train_label_coordinates.csv</code>. Those 3 that Sergio Saharovskiy posted are these missing series.</p>\n<blockquote>\n  <p>So I tried this : train_label_cords[(train_label_cords['study_id']=='3637444890') &amp; (train_label_cords['series_id']=='3892989905')] And it returned zero rows.</p>\n</blockquote>\n<p>Yes, that is what the merge is also depicting, that study_id=3637444890 series_id=3892989905 is missing in <code>train_label_coordinates.csv</code> but present in <code>train_series_descriptions.csv</code>.</p>\n<p>Edit: Also adding for this:</p>\n<blockquote>\n  <p>You should also check it by clicking on the three dots of the main folder of the competition dataset and then choosing the check for updates option.</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1303569%2F5ad26213148efa9839490b85a2007013%2FScreenshot%202024-06-30%20131955.png?generation=1719733890297553&amp;alt=media\" alt=\"No Updates Found\"></p>",
      "rawMarkdown": "The confusion that you are having is that index number `41256` in `df_train_label` and `df_train_desc.merge(df_train_label, on=['study_id', 'series_id'], how='left')` refer to 2 separate things.\nFirst on is just the 41256th row in `train_label_coordinates.csv` the other is the 41256th row in **left SQL join** of `train_series_descriptions.csv` and `train_label_coordinates.csv` on the columns `study_id` and `series_id`.\nIf you do `print(len(mrgd), len(df_train_label))` from my code cell that I gave above, you get `48695 48692`. There are 3 series missing in `train_label_coordinates.csv`. Those 3 that Sergio Saharovskiy posted are these missing series.\n\n> So I tried this : train_label_cords[(train_label_cords['study_id']=='3637444890') & (train_label_cords['series_id']=='3892989905')] And it returned zero rows.\n\nYes, that is what the merge is also depicting, that study_id=3637444890 series_id=3892989905 is missing in `train_label_coordinates.csv` but present in `train_series_descriptions.csv`.\n\nEdit: Also adding for this:\n>You should also check it by clicking on the three dots of the main folder of the competition dataset and then choosing the check for updates option.\n\n![No Updates Found](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1303569%2F5ad26213148efa9839490b85a2007013%2FScreenshot%202024-06-30%20131955.png?generation=1719733890297553&alt=media)",
      "votes": null
    },
    {
      "id": "2897087",
      "postDate": "06/30/2024 08:43:41",
      "content": "<p>Thanks a lot for this clarification, although I have not yet checked it myself. </p>",
      "rawMarkdown": "Thanks a lot for this clarification, although I have not yet checked it myself.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2890326,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/26/2024 04:29:33",
      "content": "<p>Yes, I can confirm that this is a Sagittal slice of the upper body. The cerebellum can be seen at the top and there are no lumbar vertebrae in the bottom part of the image. Thanks for pointing this case out.</p>\n<p>Fun fact: study_id 3637444890 is the only one where Sagittal T1 is used to label Spinal Canal Stenosis. All other use Sagittal T2/STIR. I guess this is the reason.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2890883,
      "author_name": "sergiosaharovskiy",
      "author_url": "",
      "post_date": "06/26/2024 11:09:20",
      "content": "<p>Yes, it is the cervical spine image, I was about to post it few days ago. Just drop it. I have not found any other incosistencies of this nature anywhere else.</p>\n<p>The algorithm your teammate has proposed is actually adds it to the input channels during the dataloading. I have not noticed a big difference in the score though with/or without it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2891500,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "06/26/2024 16:57:10",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a>. How did you identify this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2891625,
          "author_name": "sergiosaharovskiy",
          "author_url": "",
          "post_date": "06/26/2024 18:23:50",
          "content": "<p>Idk how the author  identified that and you did not ask me, though if you let me it might help answer other people questions. If you go by simple merge <code>test_series_descriptions.csv</code> and <code>train_label_coordinates.csv</code> and sort by missing coordinates - there will be some hits. I also went thru all sagittal examples, I did not find anything else like this.</p>\n<pre><code>mrgd = train_desc.merge(train_coord, on=[, ], how=)\n(mrgd[mrgd.x.isna()].to_string())\n</code></pre>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>study_id</th>\n<th>series_id</th>\n<th>series_description</th>\n<th>instance_number</th>\n<th>condition</th>\n<th>level</th>\n<th>x</th>\n<th>y</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>33978</td>\n<td>3008676218</td>\n<td>542282425</td>\n<td>Sagittal T1</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n</tr>\n<tr>\n<td>33979</td>\n<td>3008676218</td>\n<td>3636216534</td>\n<td>Axial T2</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n</tr>\n<tr>\n<td>41256</td>\n<td>3637444890</td>\n<td>3892989905</td>\n<td>Sagittal T2/STIR</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n<td>NaN</td>\n</tr>\n</tbody>\n</table>",
          "votes": null,
          "replies": [
            {
              "id": 2891730,
              "author_name": "brendanartley",
              "author_url": "",
              "post_date": "06/26/2024 20:02:12",
              "content": "<p>My fault, <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>. I should have asked you both!</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2896221,
              "author_name": "devsya",
              "author_url": "",
              "post_date": "06/29/2024 17:10:35",
              "content": "<p>Thank you for this valuable insight.  <br>\nI have noticed that total_train_label_cords_desc['instance_number'][41256] produces 5 and total_train_label_cords_desc['condition'][41256] produces 'Right Neural Foraminal Narrowing' and not nan. I am a bit confused by this output has the dataset been modified?</p>\n<p>Your responses will be highly valuable. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2896316,
                  "author_name": "coderrkj",
                  "author_url": "",
                  "post_date": "06/29/2024 18:41:34",
                  "content": "<p><a href=\"https://www.kaggle.com/devsya\" target=\"_blank\">@devsya</a>, it would be great if you could show a snippet on how you got <code>total_train_label_cords_desc</code>. <br>\nI reran the code Sergey gave in a Kaggle notebook cell:</p>\n<pre><code> pathlib  Path\n pandas  pd\n\nINPUT_DIR = Path()\ndf_train_label = pd.read_csv(INPUT_DIR / )\ndf_train_desc = pd.read_csv(INPUT_DIR / )\n\nmrgd = df_train_desc.merge(df_train_label, on=[, ], how=)\nmrgd[mrgd.x.isna()]\n</code></pre>\n<p>and the result is the same as above. To be specific to your question, <code>mrgd.loc[41256]</code> gives:</p>\n<pre><code>study_id                    3637444890\nseries_id                   3892989905\nseries_description    Sagittal T2/STIR\ninstance_number                    NaN\ncondition                          NaN\nlevel                              NaN\nx                                  NaN\ny                                  NaN\nName: 41256, dtype: object\n</code></pre>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2896799,
                      "author_name": "devsya",
                      "author_url": "",
                      "post_date": "06/30/2024 06:12:09",
                      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14118568%2Fae6104f614419dae16aa9efd654664e5%2Fscreenshot_train_label_cords.png?generation=1719727922004953&amp;alt=media\" alt=\"See This\"></p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            },
            {
              "id": 2896806,
              "author_name": "devsya",
              "author_url": "",
              "post_date": "06/30/2024 06:20:09",
              "content": "<p><a href=\"https://www.kaggle.com/coderrkj\" target=\"_blank\">@coderrkj</a> In the image you would see that 41256 points to a different series ID. So I tried this : train_label_cords[(train_label_cords['study_id']=='3637444890') &amp; (train_label_cords['series_id']=='3892989905')]<br>\nAnd it returned zero rows. I think that the dataset has been updated. You should also check it by clicking on the three dots of the main folder of the competition dataset and then choosing the check for updates option.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2896964,
                  "author_name": "coderrkj",
                  "author_url": "",
                  "post_date": "06/30/2024 07:44:58",
                  "content": "<p>The confusion that you are having is that index number <code>41256</code> in <code>df_train_label</code> and <code>df_train_desc.merge(df_train_label, on=['study_id', 'series_id'], how='left')</code> refer to 2 separate things.<br>\nFirst on is just the 41256th row in <code>train_label_coordinates.csv</code> the other is the 41256th row in <strong>left SQL join</strong> of <code>train_series_descriptions.csv</code> and <code>train_label_coordinates.csv</code> on the columns <code>study_id</code> and <code>series_id</code>.<br>\nIf you do <code>print(len(mrgd), len(df_train_label))</code> from my code cell that I gave above, you get <code>48695 48692</code>. There are 3 series missing in <code>train_label_coordinates.csv</code>. Those 3 that Sergio Saharovskiy posted are these missing series.</p>\n<blockquote>\n  <p>So I tried this : train_label_cords[(train_label_cords['study_id']=='3637444890') &amp; (train_label_cords['series_id']=='3892989905')] And it returned zero rows.</p>\n</blockquote>\n<p>Yes, that is what the merge is also depicting, that study_id=3637444890 series_id=3892989905 is missing in <code>train_label_coordinates.csv</code> but present in <code>train_series_descriptions.csv</code>.</p>\n<p>Edit: Also adding for this:</p>\n<blockquote>\n  <p>You should also check it by clicking on the three dots of the main folder of the competition dataset and then choosing the check for updates option.</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1303569%2F5ad26213148efa9839490b85a2007013%2FScreenshot%202024-06-30%20131955.png?generation=1719733890297553&amp;alt=media\" alt=\"No Updates Found\"></p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2897087,
                      "author_name": "devsya",
                      "author_url": "",
                      "post_date": "06/30/2024 08:43:41",
                      "content": "<p>Thanks a lot for this clarification, although I have not yet checked it myself. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2890264": "please take a look:    study_id: 3637444890     series_id: 3892989905\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F0f158495a8ffe4f4d2368a52d3406694%2F006.png?generation=1719372819072726&alt=media)",
    "2890326": "Yes, I can confirm that this is a Sagittal slice of the upper body. The cerebellum can be seen at the top and there are no lumbar vertebrae in the bottom part of the image. Thanks for pointing this case out.\n\nFun fact: study_id 3637444890 is the only one where Sagittal T1 is used to label Spinal Canal Stenosis. All other use Sagittal T2/STIR. I guess this is the reason.",
    "2890883": "Yes, it is the cervical spine image, I was about to post it few days ago. Just drop it. I have not found any other incosistencies of this nature anywhere else.\n\nThe algorithm your teammate has proposed is actually adds it to the input channels during the dataloading. I have not noticed a big difference in the score though with/or without it.",
    "2891500": "Thanks for sharing @lihaoweicvch. How did you identify this?",
    "2891625": "Idk how the author  identified that and you did not ask me, though if you let me it might help answer other people questions. If you go by simple merge `test_series_descriptions.csv` and `train_label_coordinates.csv` and sort by missing coordinates - there will be some hits. I also went thru all sagittal examples, I did not find anything else like this.\n```python\nmrgd = train_desc.merge(train_coord, on=['study_id', 'series_id'], how='left')\nprint(mrgd[mrgd.x.isna()].to_string())\n```\n\n|    | study_id   | series_id   | series_description | instance_number | condition | level | x   | y   |\n|----|------------|-------------|--------------------|-----------------|-----------|-------|-----|-----|\n| 33978 | 3008676218 | 542282425   | Sagittal T1         | NaN             | NaN       | NaN   | NaN | NaN |\n| 33979 | 3008676218 | 3636216534  | Axial T2            | NaN             | NaN       | NaN   | NaN | NaN |\n| 41256 | 3637444890 | 3892989905  | Sagittal T2/STIR    | NaN             | NaN       | NaN   | NaN | NaN |",
    "2891730": "My fault, @sergiosaharovskiy. I should have asked you both!",
    "2896221": "Thank you for this valuable insight.  \nI have noticed that total_train_label_cords_desc['instance_number'][41256] produces 5 and total_train_label_cords_desc['condition'][41256] produces 'Right Neural Foraminal Narrowing' and not nan. I am a bit confused by this output has the dataset been modified?\n\nYour responses will be highly valuable.",
    "2896316": "devsya, it would be great if you could show a snippet on how you got `total_train_label_cords_desc`. \nI reran the code Sergey gave in a Kaggle notebook cell:\n```python\nfrom pathlib import Path\nimport pandas as pd\n\nINPUT_DIR = Path(\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification\")\ndf_train_label = pd.read_csv(INPUT_DIR / 'train_label_coordinates.csv')\ndf_train_desc = pd.read_csv(INPUT_DIR / 'train_series_descriptions.csv')\n\nmrgd = df_train_desc.merge(df_train_label, on=['study_id', 'series_id'], how='left')\nmrgd[mrgd.x.isna()]\n```\nand the result is the same as above. To be specific to your question, `mrgd.loc[41256]` gives:\n```text\nstudy_id                    3637444890\nseries_id                   3892989905\nseries_description    Sagittal T2/STIR\ninstance_number                    NaN\ncondition                          NaN\nlevel                              NaN\nx                                  NaN\ny                                  NaN\nName: 41256, dtype: object\n```",
    "2896799": "![See This](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14118568%2Fae6104f614419dae16aa9efd654664e5%2Fscreenshot_train_label_cords.png?generation=1719727922004953&alt=media)",
    "2896806": "coderrkj In the image you would see that 41256 points to a different series ID. So I tried this : train_label_cords[(train_label_cords['study_id']=='3637444890') & (train_label_cords['series_id']=='3892989905')]\nAnd it returned zero rows. I think that the dataset has been updated. You should also check it by clicking on the three dots of the main folder of the competition dataset and then choosing the check for updates option.",
    "2896964": "The confusion that you are having is that index number `41256` in `df_train_label` and `df_train_desc.merge(df_train_label, on=['study_id', 'series_id'], how='left')` refer to 2 separate things.\nFirst on is just the 41256th row in `train_label_coordinates.csv` the other is the 41256th row in **left SQL join** of `train_series_descriptions.csv` and `train_label_coordinates.csv` on the columns `study_id` and `series_id`.\nIf you do `print(len(mrgd), len(df_train_label))` from my code cell that I gave above, you get `48695 48692`. There are 3 series missing in `train_label_coordinates.csv`. Those 3 that Sergio Saharovskiy posted are these missing series.\n\n> So I tried this : train_label_cords[(train_label_cords['study_id']=='3637444890') & (train_label_cords['series_id']=='3892989905')] And it returned zero rows.\n\nYes, that is what the merge is also depicting, that study_id=3637444890 series_id=3892989905 is missing in `train_label_coordinates.csv` but present in `train_series_descriptions.csv`.\n\nEdit: Also adding for this:\n>You should also check it by clicking on the three dots of the main folder of the competition dataset and then choosing the check for updates option.\n\n![No Updates Found](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1303569%2F5ad26213148efa9839490b85a2007013%2FScreenshot%202024-06-30%20131955.png?generation=1719733890297553&alt=media)",
    "2897087": "Thanks a lot for this clarification, although I have not yet checked it myself."
  },
  "source": "meta"
}