{
  "id": 512048,
  "title": "Suspicious wrong coordinates labels",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/512048",
  "author_name": "Hajimi Nanbeilvdou",
  "post_date": "2024-06-13T08:51:50.013000",
  "votes": 18,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi, everybody. I found some coordinates may be incorrectly annotated on images of stdudy id 4279881930 series id 3657125167 instance number 14, 15 and 16. Compared with images ahead like instance 7 and 8, the coordinate level seems get wrongly labled higher by one. (e.g. it should be l3/l4 but it is labeled as l2/l3)<br>\nI am not an expert and hope to hear any opinions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5567488%2Fc5784a6da9d135c92d1f9188af9c4f72%2F4279881930_3657125167_15.png?generation=1718268673559803&amp;alt=media\" alt=\"4279881930_3657125167_15 \"></p>",
  "messages": [
    {
      "id": 2869764,
      "postDate": "2024-06-13T08:51:50.013Z",
      "content": "<p>Hi, everybody. I found some coordinates may be incorrectly annotated on images of stdudy id 4279881930 series id 3657125167 instance number 14, 15 and 16. Compared with images ahead like instance 7 and 8, the coordinate level seems get wrongly labled higher by one. (e.g. it should be l3/l4 but it is labeled as l2/l3)<br>\nI am not an expert and hope to hear any opinions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5567488%2Fc5784a6da9d135c92d1f9188af9c4f72%2F4279881930_3657125167_15.png?generation=1718268673559803&amp;alt=media\" alt=\"4279881930_3657125167_15 \"></p>",
      "rawMarkdown": "Hi, everybody. I found some coordinates may be incorrectly annotated on images of stdudy id 4279881930 series id 3657125167 instance number 14, 15 and 16. Compared with images ahead like instance 7 and 8, the coordinate level seems get wrongly labled higher by one. (e.g. it should be l3/l4 but it is labeled as l2/l3)\nI am not an expert and hope to hear any opinions.\n![4279881930_3657125167_15 ](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5567488%2Fc5784a6da9d135c92d1f9188af9c4f72%2F4279881930_3657125167_15.png?generation=1718268673559803&alt=media)",
      "votes": 18
    },
    {
      "id": 2870939,
      "postDate": "2024-06-13T22:42:17.043Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fef9522446b7fc46900643a4861319803%2F5examples.png?generation=1718329552852857&amp;alt=media\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fef9522446b7fc46900643a4861319803%2F5examples.png?generation=1718329552852857&alt=media)",
      "votes": 5,
      "replies": [
        {
          "id": 2871177,
          "postDate": "2024-06-14T05:25:35.933Z",
          "content": "<p>Thanks for pointing this out. In addition, there seems to be a (5, 5) error in some of the slices. I have listed them here:</p>\n<table>\n<thead>\n<tr>\n<th>study_id</th>\n<th>series_id</th>\n<th>instance_number</th>\n<th>condition</th>\n<th>level</th>\n<th>x</th>\n<th>y</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>38281420</td>\n<td>880361156</td>\n<td>9</td>\n<td>Spinal Canal Stenosis</td>\n<td>L5/S1</td>\n<td>5</td>\n<td>5</td>\n</tr>\n<tr>\n<td>286903519</td>\n<td>1921917205</td>\n<td>13</td>\n<td>Spinal Canal Stenosis</td>\n<td>L5/S1</td>\n<td>5.00001</td>\n<td>5.00001</td>\n</tr>\n<tr>\n<td>665627263</td>\n<td>2231471633</td>\n<td>9</td>\n<td>Spinal Canal Stenosis</td>\n<td>L1/L2</td>\n<td>5</td>\n<td>5</td>\n</tr>\n<tr>\n<td>1438760543</td>\n<td>737753815</td>\n<td>9</td>\n<td>Spinal Canal Stenosis</td>\n<td>L1/L2</td>\n<td>5.00004</td>\n<td>4.99999</td>\n</tr>\n<tr>\n<td>1510451897</td>\n<td>1488857550</td>\n<td>9</td>\n<td>Spinal Canal Stenosis</td>\n<td>L4/L5</td>\n<td>5</td>\n<td>4.99998</td>\n</tr>\n<tr>\n<td>1880970480</td>\n<td>3736941525</td>\n<td>8</td>\n<td>Spinal Canal Stenosis</td>\n<td>L5/S1</td>\n<td>5</td>\n<td>5</td>\n</tr>\n<tr>\n<td>1901348744</td>\n<td>1490272456</td>\n<td>11</td>\n<td>Spinal Canal Stenosis</td>\n<td>L3/L4</td>\n<td>4.99995</td>\n<td>4.99997</td>\n</tr>\n<tr>\n<td>2151467507</td>\n<td>3086719329</td>\n<td>8</td>\n<td>Spinal Canal Stenosis</td>\n<td>L3/L4</td>\n<td>4.05882</td>\n<td>5</td>\n</tr>\n<tr>\n<td>2316015842</td>\n<td>1485193299</td>\n<td>13</td>\n<td>Spinal Canal Stenosis</td>\n<td>L1/L2</td>\n<td>5.00003</td>\n<td>5.00001</td>\n</tr>\n<tr>\n<td>2444340715</td>\n<td>3521409198</td>\n<td>10</td>\n<td>Spinal Canal Stenosis</td>\n<td>L5/S1</td>\n<td>5.00003</td>\n<td>4.99999</td>\n</tr>\n<tr>\n<td>2905025904</td>\n<td>816381378</td>\n<td>11</td>\n<td>Spinal Canal Stenosis</td>\n<td>L1/L2</td>\n<td>4.99998</td>\n<td>5</td>\n</tr>\n</tbody>\n</table>\n<p>These would be equivalent to the 1st (38281420, 880361156), 2nd (286903519, 1921917205) and last slice (2905025904, 816381378) above.</p>",
          "rawMarkdown": "Thanks for pointing this out. In addition, there seems to be a (5, 5) error in some of the slices. I have listed them here:\n|   study_id |   series_id |   instance_number | condition             | level   |       x |       y |\n|-----------:|------------:|------------------:|:----------------------|:--------|--------:|--------:|\n|   38281420 |   880361156 |                 9 | Spinal Canal Stenosis | L5/S1   | 5       | 5       |\n|  286903519 |  1921917205 |                13 | Spinal Canal Stenosis | L5/S1   | 5.00001 | 5.00001 |\n|  665627263 |  2231471633 |                 9 | Spinal Canal Stenosis | L1/L2   | 5       | 5       |\n| 1438760543 |   737753815 |                 9 | Spinal Canal Stenosis | L1/L2   | 5.00004 | 4.99999 |\n| 1510451897 |  1488857550 |                 9 | Spinal Canal Stenosis | L4/L5   | 5       | 4.99998 |\n| 1880970480 |  3736941525 |                 8 | Spinal Canal Stenosis | L5/S1   | 5       | 5       |\n| 1901348744 |  1490272456 |                11 | Spinal Canal Stenosis | L3/L4   | 4.99995 | 4.99997 |\n| 2151467507 |  3086719329 |                 8 | Spinal Canal Stenosis | L3/L4   | 4.05882 | 5       |\n| 2316015842 |  1485193299 |                13 | Spinal Canal Stenosis | L1/L2   | 5.00003 | 5.00001 |\n| 2444340715 |  3521409198 |                10 | Spinal Canal Stenosis | L5/S1   | 5.00003 | 4.99999 |\n| 2905025904 |   816381378 |                11 | Spinal Canal Stenosis | L1/L2   | 4.99998 | 5       |\n\nThese would be equivalent to the 1st (38281420, 880361156), 2nd (286903519, 1921917205) and last slice (2905025904, 816381378) above.",
          "votes": 6
        },
        {
          "id": 2874163,
          "postDate": "2024-06-16T04:59:12.077Z",
          "content": "<p>Thanks a lot. For pointing it out in a clear way.</p>",
          "rawMarkdown": "Thanks a lot. For pointing it out in a clear way."
        }
      ]
    },
    {
      "id": 2870341,
      "postDate": "2024-06-13T14:23:18.957Z",
      "content": "<p>I am not an expert by any means, but a side-by-side comparison of the slices reveal that at least the adjacent slices are consistent:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1303569%2Fa298e0c69d9df27fe40d7983ebb52bb2%2Fside-by-side_slices.png?generation=1718288436785505&amp;alt=media\"><br>\nHere is the code to create this:</p>\n<pre><code> pathlib  Path\n\n matplotlib.pyplot  plt\n pandas  pd\n pydicom\n\nstudy_id, series_id = , \nfig, axes = plt.subplots(, , figsize=(, ))\n\ninstance_numbers = [, , ]\n instance_number, ax  (instance_numbers, axes[]):\n    dcm_path = train_images_path / (study_id) / (series_id) / \n     dcm_path.exists()\n    dcm = pydicom.dcmread(dcm_path)\n     (dcm.InstanceNumber) == instance_number \n\n    ax.imshow(dcm.pixel_array, )\n\n    conditions = (\n        (df_train_label.study_id == study_id) &amp; (df_train_label.series_id == series_id) &amp; (df_train_label.instance_number == instance_number)\n    )\n    sub_lbls = df_train_label[conditions]\n    sub_lbls.plot(, , , ax=ax, color=)\n    ax.set_title( + .join(sub_lbls.level.values.tolist()))\n\ninstance_numbers = [, , ]\n instance_number, ax  (instance_numbers, axes[]):\n    dcm_path = train_images_path / (study_id) / (series_id) / \n     dcm_path.exists()\n    dcm = pydicom.dcmread(dcm_path)\n     (dcm.InstanceNumber) == instance_number \n\n    ax.imshow(dcm.pixel_array, )\n\n    conditions = (\n        (df_train_label.study_id == study_id) &amp; (df_train_label.series_id == series_id) &amp; (df_train_label.instance_number == instance_number)\n    )\n    sub_lbls = df_train_label[conditions]\n    sub_lbls.plot(, , , ax=ax, color=)\n    ax.set_title( + .join(sub_lbls.level.values.tolist()))\nplt.show()\n</code></pre>",
      "rawMarkdown": "I am not an expert by any means, but a side-by-side comparison of the slices reveal that at least the adjacent slices are consistent:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1303569%2Fa298e0c69d9df27fe40d7983ebb52bb2%2Fside-by-side_slices.png?generation=1718288436785505&alt=media)\nHere is the code to create this:\n```python\nfrom pathlib import Path\n\nimport matplotlib.pyplot as plt\nimport pandas as pd\nimport pydicom\n\nstudy_id, series_id = 4279881930, 3657125167\nfig, axes = plt.subplots(2, 3, figsize=(20, 10))\n\ninstance_numbers = [6, 7, 8]\nfor instance_number, ax in zip(instance_numbers, axes[0]):\n    dcm_path = train_images_path / str(study_id) / str(series_id) / f\"{instance_number}.dcm\"\n    assert dcm_path.exists()\n    dcm = pydicom.dcmread(dcm_path)\n    assert int(dcm.InstanceNumber) == instance_number # Double check\n\n    ax.imshow(dcm.pixel_array, \"gray\")\n\n    conditions = (\n        (df_train_label.study_id == study_id) & (df_train_label.series_id == series_id) & (df_train_label.instance_number == instance_number)\n    )\n    sub_lbls = df_train_label[conditions]\n    sub_lbls.plot(\"x\", \"y\", \"scatter\", ax=ax, color=\"red\")\n    ax.set_title(f\"{instance_number}:\" + \", \".join(sub_lbls.level.values.tolist()))\n    \ninstance_numbers = [14, 15, 16]\nfor instance_number, ax in zip(instance_numbers, axes[1]):\n    dcm_path = train_images_path / str(study_id) / str(series_id) / f\"{instance_number}.dcm\"\n    assert dcm_path.exists()\n    dcm = pydicom.dcmread(dcm_path)\n    assert int(dcm.InstanceNumber) == instance_number # Double check\n\n    ax.imshow(dcm.pixel_array, \"gray\")\n\n    conditions = (\n        (df_train_label.study_id == study_id) & (df_train_label.series_id == series_id) & (df_train_label.instance_number == instance_number)\n    )\n    sub_lbls = df_train_label[conditions]\n    sub_lbls.plot(\"x\", \"y\", \"scatter\", ax=ax, color=\"red\")\n    ax.set_title(f\"{instance_number}:\" + \", \".join(sub_lbls.level.values.tolist()))\nplt.show()\n```",
      "votes": 1
    },
    {
      "id": 2874166,
      "postDate": "2024-06-16T05:02:17.787Z",
      "content": "<p>Thank you. And how did you manage to find it out? Were you going through each image and checking if they are labelled correctly?</p>",
      "rawMarkdown": "Thank you. And how did you manage to find it out? Were you going through each image and checking if they are labelled correctly?",
      "replies": [
        {
          "id": 2874443,
          "postDate": "2024-06-16T09:14:25.507Z",
          "content": "<p>In my case I noticed in the examples <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> put above in comment, that 3 of the 5 images showcased points close to (5, 5). I filtered for x and y close to 5 and got the list I outlined in the reply.</p>",
          "rawMarkdown": "In my case I noticed in the examples @sergiosaharovskiy put above in comment, that 3 of the 5 images showcased points close to (5, 5). I filtered for x and y close to 5 and got the list I outlined in the reply.",
          "votes": 2,
          "replies": [
            {
              "id": 2874476,
              "postDate": "2024-06-16T10:01:18.460Z",
              "content": "<p>You mean that you applied a filter for each unique image. The filter returned the images that had multiple labels in them whose coordinates were less than 5 units (pixels?) away from each other.</p>\n<p>Do you mean this?</p>",
              "rawMarkdown": "You mean that you applied a filter for each unique image. The filter returned the images that had multiple labels in them whose coordinates were less than 5 units (pixels?) away from each other.\n\nDo you mean this?"
            },
            {
              "id": 2874480,
              "postDate": "2024-06-16T10:05:14.547Z",
              "content": "<p>Ok, now I know how you found this. Thank you</p>\n<p>The approach in my previous comment can also be very effective (in the case of sagittal views ) because correct coordinates in a single image may be at least somewhat far from each other .</p>",
              "rawMarkdown": "Ok, now I know how you found this. Thank you\n\nThe approach in my previous comment can also be very effective (in the case of sagittal views ) because correct coordinates in a single image may be at least somewhat far from each other .",
              "votes": 1
            },
            {
              "id": 2875381,
              "postDate": "2024-06-17T05:12:40.827Z",
              "content": "<p>This is the exact code for getting the above table for coordinates near (5.0, 5.0):</p>\n<pre><code> pathlib  Path\n\n pandas  pd\n\n\nINPUT_DIR = Path()\ntrain_images_path = INPUT_DIR / \n\n\ndf_train_main = pd.read_csv(INPUT_DIR / )\ndf_train_label = pd.read_csv(INPUT_DIR / )\n\n(df_train_label[((df_train_label.x - ) **  + (df_train_label.y - ) ** ) &lt;= ].to_markdown(index=))\n</code></pre>",
              "rawMarkdown": "This is the exact code for getting the above table for coordinates near (5.0, 5.0):\n```python\nfrom pathlib import Path\n\nimport pandas as pd\n\n\nINPUT_DIR = Path(\"../input/rsna-2024-lumbar-spine-degenerative-classification\")\ntrain_images_path = INPUT_DIR / \"train_images\"\n\n# read data\ndf_train_main = pd.read_csv(INPUT_DIR / 'train.csv')\ndf_train_label = pd.read_csv(INPUT_DIR / 'train_label_coordinates.csv')\n\nprint(df_train_label[((df_train_label.x - 5.0) ** 2 + (df_train_label.y - 5.0) ** 2) <= 1].to_markdown(index=False))\n```",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2871964,
      "postDate": "2024-06-14T13:46:50.923Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2870939,
      "author_name": "SSS",
      "author_url": "",
      "post_date": "2024-06-13T22:42:17.043000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fef9522446b7fc46900643a4861319803%2F5examples.png?generation=1718329552852857&amp;alt=media\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 2871177,
          "author_name": "coderRKJ",
          "author_url": "",
          "post_date": "2024-06-14T05:25:35.933000",
          "content": "<p>Thanks for pointing this out. In addition, there seems to be a (5, 5) error in some of the slices. I have listed them here:</p>\n<table>\n<thead>\n<tr>\n<th>study_id</th>\n<th>series_id</th>\n<th>instance_number</th>\n<th>condition</th>\n<th>level</th>\n<th>x</th>\n<th>y</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>38281420</td>\n<td>880361156</td>\n<td>9</td>\n<td>Spinal Canal Stenosis</td>\n<td>L5/S1</td>\n<td>5</td>\n<td>5</td>\n</tr>\n<tr>\n<td>286903519</td>\n<td>1921917205</td>\n<td>13</td>\n<td>Spinal Canal Stenosis</td>\n<td>L5/S1</td>\n<td>5.00001</td>\n<td>5.00001</td>\n</tr>\n<tr>\n<td>665627263</td>\n<td>2231471633</td>\n<td>9</td>\n<td>Spinal Canal Stenosis</td>\n<td>L1/L2</td>\n<td>5</td>\n<td>5</td>\n</tr>\n<tr>\n<td>1438760543</td>\n<td>737753815</td>\n<td>9</td>\n<td>Spinal Canal Stenosis</td>\n<td>L1/L2</td>\n<td>5.00004</td>\n<td>4.99999</td>\n</tr>\n<tr>\n<td>1510451897</td>\n<td>1488857550</td>\n<td>9</td>\n<td>Spinal Canal Stenosis</td>\n<td>L4/L5</td>\n<td>5</td>\n<td>4.99998</td>\n</tr>\n<tr>\n<td>1880970480</td>\n<td>3736941525</td>\n<td>8</td>\n<td>Spinal Canal Stenosis</td>\n<td>L5/S1</td>\n<td>5</td>\n<td>5</td>\n</tr>\n<tr>\n<td>1901348744</td>\n<td>1490272456</td>\n<td>11</td>\n<td>Spinal Canal Stenosis</td>\n<td>L3/L4</td>\n<td>4.99995</td>\n<td>4.99997</td>\n</tr>\n<tr>\n<td>2151467507</td>\n<td>3086719329</td>\n<td>8</td>\n<td>Spinal Canal Stenosis</td>\n<td>L3/L4</td>\n<td>4.05882</td>\n<td>5</td>\n</tr>\n<tr>\n<td>2316015842</td>\n<td>1485193299</td>\n<td>13</td>\n<td>Spinal Canal Stenosis</td>\n<td>L1/L2</td>\n<td>5.00003</td>\n<td>5.00001</td>\n</tr>\n<tr>\n<td>2444340715</td>\n<td>3521409198</td>\n<td>10</td>\n<td>Spinal Canal Stenosis</td>\n<td>L5/S1</td>\n<td>5.00003</td>\n<td>4.99999</td>\n</tr>\n<tr>\n<td>2905025904</td>\n<td>816381378</td>\n<td>11</td>\n<td>Spinal Canal Stenosis</td>\n<td>L1/L2</td>\n<td>4.99998</td>\n<td>5</td>\n</tr>\n</tbody>\n</table>\n<p>These would be equivalent to the 1st (38281420, 880361156), 2nd (286903519, 1921917205) and last slice (2905025904, 816381378) above.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 2874163,
          "author_name": "Devsya ",
          "author_url": "",
          "post_date": "2024-06-16T04:59:12.077000",
          "content": "<p>Thanks a lot. For pointing it out in a clear way.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2870341,
      "author_name": "coderRKJ",
      "author_url": "",
      "post_date": "2024-06-13T14:23:18.957000",
      "content": "<p>I am not an expert by any means, but a side-by-side comparison of the slices reveal that at least the adjacent slices are consistent:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1303569%2Fa298e0c69d9df27fe40d7983ebb52bb2%2Fside-by-side_slices.png?generation=1718288436785505&amp;alt=media\"><br>\nHere is the code to create this:</p>\n<pre><code> pathlib  Path\n\n matplotlib.pyplot  plt\n pandas  pd\n pydicom\n\nstudy_id, series_id = , \nfig, axes = plt.subplots(, , figsize=(, ))\n\ninstance_numbers = [, , ]\n instance_number, ax  (instance_numbers, axes[]):\n    dcm_path = train_images_path / (study_id) / (series_id) / \n     dcm_path.exists()\n    dcm = pydicom.dcmread(dcm_path)\n     (dcm.InstanceNumber) == instance_number \n\n    ax.imshow(dcm.pixel_array, )\n\n    conditions = (\n        (df_train_label.study_id == study_id) &amp; (df_train_label.series_id == series_id) &amp; (df_train_label.instance_number == instance_number)\n    )\n    sub_lbls = df_train_label[conditions]\n    sub_lbls.plot(, , , ax=ax, color=)\n    ax.set_title( + .join(sub_lbls.level.values.tolist()))\n\ninstance_numbers = [, , ]\n instance_number, ax  (instance_numbers, axes[]):\n    dcm_path = train_images_path / (study_id) / (series_id) / \n     dcm_path.exists()\n    dcm = pydicom.dcmread(dcm_path)\n     (dcm.InstanceNumber) == instance_number \n\n    ax.imshow(dcm.pixel_array, )\n\n    conditions = (\n        (df_train_label.study_id == study_id) &amp; (df_train_label.series_id == series_id) &amp; (df_train_label.instance_number == instance_number)\n    )\n    sub_lbls = df_train_label[conditions]\n    sub_lbls.plot(, , , ax=ax, color=)\n    ax.set_title( + .join(sub_lbls.level.values.tolist()))\nplt.show()\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2874166,
      "author_name": "Devsya ",
      "author_url": "",
      "post_date": "2024-06-16T05:02:17.787000",
      "content": "<p>Thank you. And how did you manage to find it out? Were you going through each image and checking if they are labelled correctly?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2874443,
          "author_name": "coderRKJ",
          "author_url": "",
          "post_date": "2024-06-16T09:14:25.507000",
          "content": "<p>In my case I noticed in the examples <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> put above in comment, that 3 of the 5 images showcased points close to (5, 5). I filtered for x and y close to 5 and got the list I outlined in the reply.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2874476,
              "author_name": "Devsya ",
              "author_url": "",
              "post_date": "2024-06-16T10:01:18.460000",
              "content": "<p>You mean that you applied a filter for each unique image. The filter returned the images that had multiple labels in them whose coordinates were less than 5 units (pixels?) away from each other.</p>\n<p>Do you mean this?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2874480,
              "author_name": "Devsya ",
              "author_url": "",
              "post_date": "2024-06-16T10:05:14.547000",
              "content": "<p>Ok, now I know how you found this. Thank you</p>\n<p>The approach in my previous comment can also be very effective (in the case of sagittal views ) because correct coordinates in a single image may be at least somewhat far from each other .</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2875381,
              "author_name": "coderRKJ",
              "author_url": "",
              "post_date": "2024-06-17T05:12:40.827000",
              "content": "<p>This is the exact code for getting the above table for coordinates near (5.0, 5.0):</p>\n<pre><code> pathlib  Path\n\n pandas  pd\n\n\nINPUT_DIR = Path()\ntrain_images_path = INPUT_DIR / \n\n\ndf_train_main = pd.read_csv(INPUT_DIR / )\ndf_train_label = pd.read_csv(INPUT_DIR / )\n\n(df_train_label[((df_train_label.x - ) **  + (df_train_label.y - ) ** ) &lt;= ].to_markdown(index=))\n</code></pre>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2871964,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-14T13:46:50.923000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2869764": "Hi, everybody. I found some coordinates may be incorrectly annotated on images of stdudy id 4279881930 series id 3657125167 instance number 14, 15 and 16. Compared with images ahead like instance 7 and 8, the coordinate level seems get wrongly labled higher by one. (e.g. it should be l3/l4 but it is labeled as l2/l3)\nI am not an expert and hope to hear any opinions.\n![4279881930_3657125167_15 ](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5567488%2Fc5784a6da9d135c92d1f9188af9c4f72%2F4279881930_3657125167_15.png?generation=1718268673559803&alt=media)",
    "2870939": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fef9522446b7fc46900643a4861319803%2F5examples.png?generation=1718329552852857&alt=media)",
    "2870341": "I am not an expert by any means, but a side-by-side comparison of the slices reveal that at least the adjacent slices are consistent:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1303569%2Fa298e0c69d9df27fe40d7983ebb52bb2%2Fside-by-side_slices.png?generation=1718288436785505&alt=media)\nHere is the code to create this:\n```python\nfrom pathlib import Path\n\nimport matplotlib.pyplot as plt\nimport pandas as pd\nimport pydicom\n\nstudy_id, series_id = 4279881930, 3657125167\nfig, axes = plt.subplots(2, 3, figsize=(20, 10))\n\ninstance_numbers = [6, 7, 8]\nfor instance_number, ax in zip(instance_numbers, axes[0]):\n    dcm_path = train_images_path / str(study_id) / str(series_id) / f\"{instance_number}.dcm\"\n    assert dcm_path.exists()\n    dcm = pydicom.dcmread(dcm_path)\n    assert int(dcm.InstanceNumber) == instance_number # Double check\n\n    ax.imshow(dcm.pixel_array, \"gray\")\n\n    conditions = (\n        (df_train_label.study_id == study_id) & (df_train_label.series_id == series_id) & (df_train_label.instance_number == instance_number)\n    )\n    sub_lbls = df_train_label[conditions]\n    sub_lbls.plot(\"x\", \"y\", \"scatter\", ax=ax, color=\"red\")\n    ax.set_title(f\"{instance_number}:\" + \", \".join(sub_lbls.level.values.tolist()))\n    \ninstance_numbers = [14, 15, 16]\nfor instance_number, ax in zip(instance_numbers, axes[1]):\n    dcm_path = train_images_path / str(study_id) / str(series_id) / f\"{instance_number}.dcm\"\n    assert dcm_path.exists()\n    dcm = pydicom.dcmread(dcm_path)\n    assert int(dcm.InstanceNumber) == instance_number # Double check\n\n    ax.imshow(dcm.pixel_array, \"gray\")\n\n    conditions = (\n        (df_train_label.study_id == study_id) & (df_train_label.series_id == series_id) & (df_train_label.instance_number == instance_number)\n    )\n    sub_lbls = df_train_label[conditions]\n    sub_lbls.plot(\"x\", \"y\", \"scatter\", ax=ax, color=\"red\")\n    ax.set_title(f\"{instance_number}:\" + \", \".join(sub_lbls.level.values.tolist()))\nplt.show()\n```",
    "2874166": "Thank you. And how did you manage to find it out? Were you going through each image and checking if they are labelled correctly?",
    "2871964": ""
  }
}