{
  "id": 514887,
  "title": "Questions about the Dataset",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/514887",
  "author_name": "",
  "post_date": "2024-06-26T03:28:56.358293100Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p><strong>Note</strong>: before stating my questions I want to mention that English is not my native language. I'm helping myself with a translator. So, in advance, I apologize if there are any syntactic or semantic errors in the writing.</p>\n<p><strong>1.</strong> It is known that the studies contain between 2 and 6 series and that 82.6% of the studies contain 3 series. As I was reading and reviewing, those studies that have more than 3 series have more than one \"Axial T2\" series (mostly). What is the approach you are following for studies containing more than 3 series? Will it be feasible to combine the excess series into a single series?</p>\n<p><strong>2.</strong> The 'train_label_coordinates' file has a total of 48692 rows, where there are cases where different rows refer to the same image (instance_number) out of a total of 147218 listed in the 'train_images' directory, so they are not using the 100% images in training data. Is this interpretation correct? If so, the challenge is to give use to those discarded images in the csv files of the training set?</p>\n<p>I greatly appreciate any clarification and contribution regarding these doubts!</p>",
  "messages": [
    {
      "id": "2890256",
      "postDate": "06/26/2024 03:28:56",
      "content": "<p><strong>Note</strong>: before stating my questions I want to mention that English is not my native language. I'm helping myself with a translator. So, in advance, I apologize if there are any syntactic or semantic errors in the writing.</p>\n<p><strong>1.</strong> It is known that the studies contain between 2 and 6 series and that 82.6% of the studies contain 3 series. As I was reading and reviewing, those studies that have more than 3 series have more than one \"Axial T2\" series (mostly). What is the approach you are following for studies containing more than 3 series? Will it be feasible to combine the excess series into a single series?</p>\n<p><strong>2.</strong> The 'train_label_coordinates' file has a total of 48692 rows, where there are cases where different rows refer to the same image (instance_number) out of a total of 147218 listed in the 'train_images' directory, so they are not using the 100% images in training data. Is this interpretation correct? If so, the challenge is to give use to those discarded images in the csv files of the training set?</p>\n<p>I greatly appreciate any clarification and contribution regarding these doubts!</p>",
      "rawMarkdown": "**Note**: before stating my questions I want to mention that English is not my native language. I'm helping myself with a translator. So, in advance, I apologize if there are any syntactic or semantic errors in the writing.\n\n**1.** It is known that the studies contain between 2 and 6 series and that 82.6% of the studies contain 3 series. As I was reading and reviewing, those studies that have more than 3 series have more than one \"Axial T2\" series (mostly). What is the approach you are following for studies containing more than 3 series? Will it be feasible to combine the excess series into a single series?\n\n**2.** The 'train_label_coordinates' file has a total of 48692 rows, where there are cases where different rows refer to the same image (instance_number) out of a total of 147218 listed in the 'train_images' directory, so they are not using the 100% images in training data. Is this interpretation correct? If so, the challenge is to give use to those discarded images in the csv files of the training set?\n\nI greatly appreciate any clarification and contribution regarding these doubts!",
      "votes": null
    },
    {
      "id": "2890367",
      "postDate": "06/26/2024 04:53:43",
      "content": "<p>See the following observation:</p>\n<table>\n<thead>\n<tr>\n<th>condition</th>\n<th>Axial T2</th>\n<th>Sagittal T1</th>\n<th>Sagittal T2/STIR</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Left Subarticular Stenosis</td>\n<td>9608</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Right Subarticular Stenosis</td>\n<td>9612</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Left Neural Foraminal Narrowing</td>\n<td>0</td>\n<td>9860</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Right Neural Foraminal Narrowing</td>\n<td>0</td>\n<td>9859</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Spinal Canal Stenosis</td>\n<td>0</td>\n<td>5</td>\n<td>9748</td>\n</tr>\n</tbody>\n</table>\n<p>Each condition has a particular series to refer to. As for why there is a deviation with 5 Spinal Canal Stenosis with Sagittal T1 see my comment <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/514891#2890326\" target=\"_blank\">here</a>.</p>\n<p>Let me start with question 1) In the case of duplicate Sagittal T1 series, it was due to there being on for the left and another for the right. As for the axial ones, it is too irregular for me to completely tell. But different axial series contain different levels and slightly different orientations also. See the post: <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/507101\" target=\"_blank\">Radiologist's Insights I</a></p>\n<p>Most starter notebooks I have seen seem to uniformly pick slices from the all the available slices. Some pick a fixed number of slices from each description (Axial T2, Sagittal T1 and Sagittal T2/STIR)</p>\n<p>2) <code>train_label_coordinates.csv</code> if having multiple coordinates in a single slice, it is because of it being a sagittal slice with different levels or an Axial T2 series with both Left and Right Subarticular Stenosis of the same level labeled. Here by level, I mean L1/L2 to L5/S1.</p>",
      "rawMarkdown": "See the following observation:\n| condition                        |   Axial T2 |   Sagittal T1 |   Sagittal T2/STIR |\n|:---------------------------------|-----------:|--------------:|-------------------:|\n| Left Subarticular Stenosis       |       9608 |             0 |                  0 |\n| Right Subarticular Stenosis      |       9612 |             0 |                  0 |\n| Left Neural Foraminal Narrowing  |          0 |          9860 |                  0 |\n| Right Neural Foraminal Narrowing |          0 |          9859 |                  0 |\n| Spinal Canal Stenosis            |          0 |             5 |               9748 |\n\nEach condition has a particular series to refer to. As for why there is a deviation with 5 Spinal Canal Stenosis with Sagittal T1 see my comment [here](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/514891#2890326).\n\nLet me start with question 1) In the case of duplicate Sagittal T1 series, it was due to there being on for the left and another for the right. As for the axial ones, it is too irregular for me to completely tell. But different axial series contain different levels and slightly different orientations also. See the post: [Radiologist's Insights I](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/507101)\n\nMost starter notebooks I have seen seem to uniformly pick slices from the all the available slices. Some pick a fixed number of slices from each description (Axial T2, Sagittal T1 and Sagittal T2/STIR)\n\n2) `train_label_coordinates.csv` if having multiple coordinates in a single slice, it is because of it being a sagittal slice with different levels or an Axial T2 series with both Left and Right Subarticular Stenosis of the same level labeled. Here by level, I mean L1/L2 to L5/S1.",
      "votes": null
    },
    {
      "id": "2891984",
      "postDate": "06/27/2024 03:25:55",
      "content": "<p>Regarding question 2) I understand that the same sagittal section is used to evaluate the different levels. In that case, the number of unique images in train_label_coordinates.csv would be much less than 48692. So, rephrasing my question: for the creation of the train_label_coordinates.csv file and the definition of its labels it uses only part of the images of each study?</p>\n<p>Thanks for your observations. I really appreciate your response. It's really useful to me.</p>",
      "rawMarkdown": "Regarding question 2) I understand that the same sagittal section is used to evaluate the different levels. In that case, the number of unique images in train_label_coordinates.csv would be much less than 48692. So, rephrasing my question: for the creation of the train_label_coordinates.csv file and the definition of its labels it uses only part of the images of each study?\n\nThanks for your observations. I really appreciate your response. It's really useful to me.",
      "votes": null
    },
    {
      "id": "2892185",
      "postDate": "06/27/2024 06:03:13",
      "content": "<p>Yes, your interpretation is correct that 100% of the dataset would not be involved in classification of the 25 condition-level combinations. One of the tasks that we need to do, while prediction, is to find the region of interest (ROI) for each level and condition combo. This will leave other regions not in the 25 ROIs unused. </p>",
      "rawMarkdown": "Yes, your interpretation is correct that 100% of the dataset would not be involved in classification of the 25 condition-level combinations. One of the tasks that we need to do, while prediction, is to find the region of interest (ROI) for each level and condition combo. This will leave other regions not in the 25 ROIs unused.",
      "votes": null
    },
    {
      "id": "2897510",
      "postDate": "06/30/2024 14:44:23",
      "content": "<p><a href=\"https://www.kaggle.com/coderrkj\" target=\"_blank\">@coderrkj</a>  thanks~!</p>",
      "rawMarkdown": "coderrkj  thanks~!",
      "votes": null
    },
    {
      "id": "2914500",
      "postDate": "07/10/2024 01:49:39",
      "content": "<p>Thank you so much!</p>",
      "rawMarkdown": "Thank you so much!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2890367,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/26/2024 04:53:43",
      "content": "<p>See the following observation:</p>\n<table>\n<thead>\n<tr>\n<th>condition</th>\n<th>Axial T2</th>\n<th>Sagittal T1</th>\n<th>Sagittal T2/STIR</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Left Subarticular Stenosis</td>\n<td>9608</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Right Subarticular Stenosis</td>\n<td>9612</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Left Neural Foraminal Narrowing</td>\n<td>0</td>\n<td>9860</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Right Neural Foraminal Narrowing</td>\n<td>0</td>\n<td>9859</td>\n<td>0</td>\n</tr>\n<tr>\n<td>Spinal Canal Stenosis</td>\n<td>0</td>\n<td>5</td>\n<td>9748</td>\n</tr>\n</tbody>\n</table>\n<p>Each condition has a particular series to refer to. As for why there is a deviation with 5 Spinal Canal Stenosis with Sagittal T1 see my comment <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/514891#2890326\" target=\"_blank\">here</a>.</p>\n<p>Let me start with question 1) In the case of duplicate Sagittal T1 series, it was due to there being on for the left and another for the right. As for the axial ones, it is too irregular for me to completely tell. But different axial series contain different levels and slightly different orientations also. See the post: <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/507101\" target=\"_blank\">Radiologist's Insights I</a></p>\n<p>Most starter notebooks I have seen seem to uniformly pick slices from the all the available slices. Some pick a fixed number of slices from each description (Axial T2, Sagittal T1 and Sagittal T2/STIR)</p>\n<p>2) <code>train_label_coordinates.csv</code> if having multiple coordinates in a single slice, it is because of it being a sagittal slice with different levels or an Axial T2 series with both Left and Right Subarticular Stenosis of the same level labeled. Here by level, I mean L1/L2 to L5/S1.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2891984,
          "author_name": "alvarosaiquita",
          "author_url": "",
          "post_date": "06/27/2024 03:25:55",
          "content": "<p>Regarding question 2) I understand that the same sagittal section is used to evaluate the different levels. In that case, the number of unique images in train_label_coordinates.csv would be much less than 48692. So, rephrasing my question: for the creation of the train_label_coordinates.csv file and the definition of its labels it uses only part of the images of each study?</p>\n<p>Thanks for your observations. I really appreciate your response. It's really useful to me.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2892185,
              "author_name": "coderrkj",
              "author_url": "",
              "post_date": "06/27/2024 06:03:13",
              "content": "<p>Yes, your interpretation is correct that 100% of the dataset would not be involved in classification of the 25 condition-level combinations. One of the tasks that we need to do, while prediction, is to find the region of interest (ROI) for each level and condition combo. This will leave other regions not in the 25 ROIs unused. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2897510,
                  "author_name": "christina0626",
                  "author_url": "",
                  "post_date": "06/30/2024 14:44:23",
                  "content": "<p><a href=\"https://www.kaggle.com/coderrkj\" target=\"_blank\">@coderrkj</a>  thanks~!</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2914500,
                  "author_name": "alvarosaiquita",
                  "author_url": "",
                  "post_date": "07/10/2024 01:49:39",
                  "content": "<p>Thank you so much!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2890256": "**Note**: before stating my questions I want to mention that English is not my native language. I'm helping myself with a translator. So, in advance, I apologize if there are any syntactic or semantic errors in the writing.\n\n**1.** It is known that the studies contain between 2 and 6 series and that 82.6% of the studies contain 3 series. As I was reading and reviewing, those studies that have more than 3 series have more than one \"Axial T2\" series (mostly). What is the approach you are following for studies containing more than 3 series? Will it be feasible to combine the excess series into a single series?\n\n**2.** The 'train_label_coordinates' file has a total of 48692 rows, where there are cases where different rows refer to the same image (instance_number) out of a total of 147218 listed in the 'train_images' directory, so they are not using the 100% images in training data. Is this interpretation correct? If so, the challenge is to give use to those discarded images in the csv files of the training set?\n\nI greatly appreciate any clarification and contribution regarding these doubts!",
    "2890367": "See the following observation:\n| condition                        |   Axial T2 |   Sagittal T1 |   Sagittal T2/STIR |\n|:---------------------------------|-----------:|--------------:|-------------------:|\n| Left Subarticular Stenosis       |       9608 |             0 |                  0 |\n| Right Subarticular Stenosis      |       9612 |             0 |                  0 |\n| Left Neural Foraminal Narrowing  |          0 |          9860 |                  0 |\n| Right Neural Foraminal Narrowing |          0 |          9859 |                  0 |\n| Spinal Canal Stenosis            |          0 |             5 |               9748 |\n\nEach condition has a particular series to refer to. As for why there is a deviation with 5 Spinal Canal Stenosis with Sagittal T1 see my comment [here](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/514891#2890326).\n\nLet me start with question 1) In the case of duplicate Sagittal T1 series, it was due to there being on for the left and another for the right. As for the axial ones, it is too irregular for me to completely tell. But different axial series contain different levels and slightly different orientations also. See the post: [Radiologist's Insights I](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/507101)\n\nMost starter notebooks I have seen seem to uniformly pick slices from the all the available slices. Some pick a fixed number of slices from each description (Axial T2, Sagittal T1 and Sagittal T2/STIR)\n\n2) `train_label_coordinates.csv` if having multiple coordinates in a single slice, it is because of it being a sagittal slice with different levels or an Axial T2 series with both Left and Right Subarticular Stenosis of the same level labeled. Here by level, I mean L1/L2 to L5/S1.",
    "2891984": "Regarding question 2) I understand that the same sagittal section is used to evaluate the different levels. In that case, the number of unique images in train_label_coordinates.csv would be much less than 48692. So, rephrasing my question: for the creation of the train_label_coordinates.csv file and the definition of its labels it uses only part of the images of each study?\n\nThanks for your observations. I really appreciate your response. It's really useful to me.",
    "2892185": "Yes, your interpretation is correct that 100% of the dataset would not be involved in classification of the 25 condition-level combinations. One of the tasks that we need to do, while prediction, is to find the region of interest (ROI) for each level and condition combo. This will leave other regions not in the 25 ROIs unused.",
    "2897510": "coderrkj  thanks~!",
    "2914500": "Thank you so much!"
  },
  "source": "meta"
}