{
  "id": 537510,
  "title": "Creating submission.csv",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/537510",
  "author_name": "",
  "post_date": "2024-10-03T14:44:49.006646700Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Can Someone help me with creating submission.csv.</p>\n<p>For each series_id in test_images<br>\nand for each image in series_id<br>\nwe've to predict (5 conditions * 5 levels = 25) labels with normal, moderate and severe level.</p>\n<p>so submission format shouldn't be like this<br>\nrow_id, series_id, conditions, normal, moderate and severe.</p>",
  "messages": [
    {
      "id": "3005951",
      "postDate": "10/03/2024 14:44:49",
      "content": "<p>Can Someone help me with creating submission.csv.</p>\n<p>For each series_id in test_images<br>\nand for each image in series_id<br>\nwe've to predict (5 conditions * 5 levels = 25) labels with normal, moderate and severe level.</p>\n<p>so submission format shouldn't be like this<br>\nrow_id, series_id, conditions, normal, moderate and severe.</p>",
      "rawMarkdown": "Can Someone help me with creating submission.csv.\n\nFor each series_id in test_images\nand for each image in series_id\nwe've to predict (5 conditions * 5 levels = 25) labels with normal, moderate and severe level.\n\nso submission format shouldn't be like this\nrow_id, series_id, conditions, normal, moderate and severe.",
      "votes": null
    },
    {
      "id": "3006505",
      "postDate": "10/04/2024 08:48:03",
      "content": "<p>So we have to predict severity of each condition in each level. There sum up to 3 severity * 5 conditions * 5 levels = 75 that for 1 study (or 1 person/patient). Series is just different MRI planes of the same person. <code>row_id</code> is already included condition. Therefore, <code>row_id</code> is already enough. You can check out <code>sample_submission.csv</code>.</p>",
      "rawMarkdown": "So we have to predict severity of each condition in each level. There sum up to 3 severity * 5 conditions * 5 levels = 75 that for 1 study (or 1 person/patient). Series is just different MRI planes of the same person. `row_id` is already included condition. Therefore, `row_id` is already enough. You can check out `sample_submission.csv`.",
      "votes": null
    },
    {
      "id": "3006696",
      "postDate": "10/04/2024 12:47:29",
      "content": "<p>Thanks. But so far what I understood is that for each images in a study_id has share same labels associated with that study_id in train.csv. Is that correct? Can u please take a look on my code for x_train and y_train. </p>\n<p>And for submission.csv I still couldn't understand. Can u please further eleborate.</p>\n<p>Here's my code:<br>\n**X_train = []<br>\ny_train = []</p>\n<p>base_train_dir = 'train_images'</p>\n<p>def get_image_path(study_id, series_id, instance_number, base_dir):<br>\n    return path + f\"{base_dir}/{study_id}/{series_id}/{instance_number}.dcm\"</p>\n<p>def get_one_hot_labels(study_id):</p>\n<pre><code>row = df_train == study_id]\n\n not row:\n\n    labels = row\n    one_hot_labels = \n\n       labels:\n\n          == :\n            one_hot_labels()\n        elif  == :\n            one_hot_labels()\n        elif  == :\n            one_hot_labels()\n        :\n            one_hot_labels()\n    return one_hot_labels\n:\n    return None\n</code></pre>\n<p>for _, row in df_train_label.iterrows():<br>\n    study_id = row['study_id']<br>\n    series_id = row['series_id']<br>\n    instance_number = row['instance_number']</p>\n<pre><code>image_path = get_image_path(study_id, series_id, instance_number, base_train_dir)\n\nds = pydicom.dcmread(image_path)\nimage_data = ds.pixel_array\nimage_data = image_data.astype(.float32) / .(image_data)\n\nimage_data = cv2.resize(image_data, (, ))\n\nX_train.(image_data)\n\n = get_one_hot_labels(study_id)\n    None:\n    y_train.()**\n</code></pre>",
      "rawMarkdown": "Thanks. But so far what I understood is that for each images in a study_id has share same labels associated with that study_id in train.csv. Is that correct? Can u please take a look on my code for x_train and y_train. \n\nAnd for submission.csv I still couldn't understand. Can u please further eleborate.\n\nHere's my code:\n**X_train = []\ny_train = []\n\nbase_train_dir = 'train_images'\n\ndef get_image_path(study_id, series_id, instance_number, base_dir):\n    return path + f\"{base_dir}/{study_id}/{series_id}/{instance_number}.dcm\"\n\ndef get_one_hot_labels(study_id):\n    \n    row = df_train[df_train['study_id'] == study_id]\n    \n    if not row.empty:\n        \n        labels = row.iloc[0, 1:].values\n        one_hot_labels = []\n        \n        for label in labels:\n            \n            if label == \"Normal/Mild\":\n                one_hot_labels.extend([1, 0, 0])\n            elif label == 'Moderate':\n                one_hot_labels.extend([0, 1, 0])\n            elif label == 'Severe':\n                one_hot_labels.extend([0, 0, 1])\n            else:\n                one_hot_labels.extend([0, 0, 0])\n        return one_hot_labels\n    else:\n        return None\n    \nfor _, row in df_train_label.iterrows():\n    study_id = row['study_id']\n    series_id = row['series_id']\n    instance_number = row['instance_number']\n    \n    image_path = get_image_path(study_id, series_id, instance_number, base_train_dir)\n    \n    ds = pydicom.dcmread(image_path)\n    image_data = ds.pixel_array\n    image_data = image_data.astype(np.float32) / np.max(image_data)\n    \n    image_data = cv2.resize(image_data, (224, 224))\n\n    X_train.append(image_data)\n    \n    labels = get_one_hot_labels(study_id)\n    if labels is not None:\n        y_train.append(labels)**",
      "votes": null
    },
    {
      "id": "3006742",
      "postDate": "10/04/2024 13:59:30",
      "content": "<blockquote>\n  <p>for each images in a study_id has share same labels associated with that study_id in train.csv. Is that correct?</p>\n</blockquote>\n<p>Yes, you are correct. </p>\n<p>How you converting your labels is also depend on how you define you model. They seems fine if your CrossEntropy target is class probabilities. I recommend you study <a href=\"https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline\" target=\"_blank\">this very well done notebook</a> especially in Dataset and train loop section. The author also demonstated dataloader which output class indices rather than class probabilities style. It might take time to understand but I think it is very time worthy going forward.</p>",
      "rawMarkdown": ">for each images in a study_id has share same labels associated with that study_id in train.csv. Is that correct?\n\nYes, you are correct. \n\nHow you converting your labels is also depend on how you define you model. They seems fine if your CrossEntropy target is class probabilities. I recommend you study [this very well done notebook](https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline) especially in Dataset and train loop section. The author also demonstated dataloader which output class indices rather than class probabilities style. It might take time to understand but I think it is very time worthy going forward.",
      "votes": null
    },
    {
      "id": "3006763",
      "postDate": "10/04/2024 14:25:55",
      "content": "<p>Thanks, I'll look into it. But Please help me if I need more help.</p>",
      "rawMarkdown": "Thanks, I'll look into it. But Please help me if I need more help.",
      "votes": null
    },
    {
      "id": "3007066",
      "postDate": "10/04/2024 22:01:04",
      "content": "<p>I've looked into it. It's a bit complicated for me. I've my own understanding… Please correct me.</p>\n<p>for testing…<br>\nlets assume that we've xyz study_id and 3 associated series_ids. Each series_id has different number of images. Lets say we've total of 20 images (from 3 series_id) and for each image I've to make prediction for 75 labels (in my case). So predictions array should've shape (20, 75). Correct?</p>\n<p>Then how to align this predictions array into form like submission.csv. Can u help me prepare submission.csv with this approach.</p>\n<p>Can u tell If I had chosen one of the correct solutions.</p>",
      "rawMarkdown": "I've looked into it. It's a bit complicated for me. I've my own understanding... Please correct me.\n\nfor testing...\nlets assume that we've xyz study_id and 3 associated series_ids. Each series_id has different number of images. Lets say we've total of 20 images (from 3 series_id) and for each image I've to make prediction for 75 labels (in my case). So predictions array should've shape (20, 75). Correct?\n\nThen how to align this predictions array into form like submission.csv. Can u help me prepare submission.csv with this approach.\n\nCan u tell If I had chosen one of the correct solutions.",
      "votes": null
    },
    {
      "id": "3007236",
      "postDate": "10/05/2024 06:18:25",
      "content": "<p>Most common way to do is stack images (1 image as 1 channel). For example, in the notebook, Author chose 10 for each series. So the final shape in 1 batch for individual study is [1, 30, 512, 512] (batch_size, channels, img_size, img_size). The output shape should be (batch_size, 75) in your case. 1 batch for 1 study. This is might be the most straight forward and easier solution. However, there is no correct solution. You need to try different approach for youself. </p>",
      "rawMarkdown": "Most common way to do is stack images (1 image as 1 channel). For example, in the notebook, Author chose 10 for each series. So the final shape in 1 batch for individual study is [1, 30, 512, 512] (batch_size, channels, img_size, img_size). The output shape should be (batch_size, 75) in your case. 1 batch for 1 study. This is might be the most straight forward and easier solution. However, there is no correct solution. You need to try different approach for youself.",
      "votes": null
    },
    {
      "id": "3008582",
      "postDate": "10/06/2024 19:14:18",
      "content": "<p>im kinda having some trouble trying to figure the submission format as well… </p>\n<p>first, how ima gonna associate each image to their respective condition+level if theres no such information on the csv file. they only refer to the series name. there are multiple images inside each series folder in the test, but each condition has 2 sides (left/right), except for spinal canal (sagittal t2/stir) and 5 levels! i know that certain image inside the subfolder is refered to a certain condition, but what side and what level? </p>\n<p>second, the fact there are multiple images on the test set also confuses me a bit… should i predict the probabilities for the severity scores( mild/moderate/severe) for all images then kinda average them?</p>\n<p>thank you for any clarifications! good luck yall.</p>",
      "rawMarkdown": "im kinda having some trouble trying to figure the submission format as well... \n\nfirst, how ima gonna associate each image to their respective condition+level if theres no such information on the csv file. they only refer to the series name. there are multiple images inside each series folder in the test, but each condition has 2 sides (left/right), except for spinal canal (sagittal t2/stir) and 5 levels! i know that certain image inside the subfolder is refered to a certain condition, but what side and what level? \n\nsecond, the fact there are multiple images on the test set also confuses me a bit... should i predict the probabilities for the severity scores( mild/moderate/severe) for all images then kinda average them?\n\nthank you for any clarifications! good luck yall.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3006505,
      "author_name": "sunpnwt12",
      "author_url": "",
      "post_date": "10/04/2024 08:48:03",
      "content": "<p>So we have to predict severity of each condition in each level. There sum up to 3 severity * 5 conditions * 5 levels = 75 that for 1 study (or 1 person/patient). Series is just different MRI planes of the same person. <code>row_id</code> is already included condition. Therefore, <code>row_id</code> is already enough. You can check out <code>sample_submission.csv</code>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3006696,
          "author_name": "ammughal",
          "author_url": "",
          "post_date": "10/04/2024 12:47:29",
          "content": "<p>Thanks. But so far what I understood is that for each images in a study_id has share same labels associated with that study_id in train.csv. Is that correct? Can u please take a look on my code for x_train and y_train. </p>\n<p>And for submission.csv I still couldn't understand. Can u please further eleborate.</p>\n<p>Here's my code:<br>\n**X_train = []<br>\ny_train = []</p>\n<p>base_train_dir = 'train_images'</p>\n<p>def get_image_path(study_id, series_id, instance_number, base_dir):<br>\n    return path + f\"{base_dir}/{study_id}/{series_id}/{instance_number}.dcm\"</p>\n<p>def get_one_hot_labels(study_id):</p>\n<pre><code>row = df_train == study_id]\n\n not row:\n\n    labels = row\n    one_hot_labels = \n\n       labels:\n\n          == :\n            one_hot_labels()\n        elif  == :\n            one_hot_labels()\n        elif  == :\n            one_hot_labels()\n        :\n            one_hot_labels()\n    return one_hot_labels\n:\n    return None\n</code></pre>\n<p>for _, row in df_train_label.iterrows():<br>\n    study_id = row['study_id']<br>\n    series_id = row['series_id']<br>\n    instance_number = row['instance_number']</p>\n<pre><code>image_path = get_image_path(study_id, series_id, instance_number, base_train_dir)\n\nds = pydicom.dcmread(image_path)\nimage_data = ds.pixel_array\nimage_data = image_data.astype(.float32) / .(image_data)\n\nimage_data = cv2.resize(image_data, (, ))\n\nX_train.(image_data)\n\n = get_one_hot_labels(study_id)\n    None:\n    y_train.()**\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 3006742,
              "author_name": "sunpnwt12",
              "author_url": "",
              "post_date": "10/04/2024 13:59:30",
              "content": "<blockquote>\n  <p>for each images in a study_id has share same labels associated with that study_id in train.csv. Is that correct?</p>\n</blockquote>\n<p>Yes, you are correct. </p>\n<p>How you converting your labels is also depend on how you define you model. They seems fine if your CrossEntropy target is class probabilities. I recommend you study <a href=\"https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline\" target=\"_blank\">this very well done notebook</a> especially in Dataset and train loop section. The author also demonstated dataloader which output class indices rather than class probabilities style. It might take time to understand but I think it is very time worthy going forward.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3006763,
                  "author_name": "ammughal",
                  "author_url": "",
                  "post_date": "10/04/2024 14:25:55",
                  "content": "<p>Thanks, I'll look into it. But Please help me if I need more help.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3007066,
                  "author_name": "ammughal",
                  "author_url": "",
                  "post_date": "10/04/2024 22:01:04",
                  "content": "<p>I've looked into it. It's a bit complicated for me. I've my own understanding… Please correct me.</p>\n<p>for testing…<br>\nlets assume that we've xyz study_id and 3 associated series_ids. Each series_id has different number of images. Lets say we've total of 20 images (from 3 series_id) and for each image I've to make prediction for 75 labels (in my case). So predictions array should've shape (20, 75). Correct?</p>\n<p>Then how to align this predictions array into form like submission.csv. Can u help me prepare submission.csv with this approach.</p>\n<p>Can u tell If I had chosen one of the correct solutions.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3007236,
                      "author_name": "sunpnwt12",
                      "author_url": "",
                      "post_date": "10/05/2024 06:18:25",
                      "content": "<p>Most common way to do is stack images (1 image as 1 channel). For example, in the notebook, Author chose 10 for each series. So the final shape in 1 batch for individual study is [1, 30, 512, 512] (batch_size, channels, img_size, img_size). The output shape should be (batch_size, 75) in your case. 1 batch for 1 study. This is might be the most straight forward and easier solution. However, there is no correct solution. You need to try different approach for youself. </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3008582,
                          "author_name": "evandrocardozo",
                          "author_url": "",
                          "post_date": "10/06/2024 19:14:18",
                          "content": "<p>im kinda having some trouble trying to figure the submission format as well… </p>\n<p>first, how ima gonna associate each image to their respective condition+level if theres no such information on the csv file. they only refer to the series name. there are multiple images inside each series folder in the test, but each condition has 2 sides (left/right), except for spinal canal (sagittal t2/stir) and 5 levels! i know that certain image inside the subfolder is refered to a certain condition, but what side and what level? </p>\n<p>second, the fact there are multiple images on the test set also confuses me a bit… should i predict the probabilities for the severity scores( mild/moderate/severe) for all images then kinda average them?</p>\n<p>thank you for any clarifications! good luck yall.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3005951": "Can Someone help me with creating submission.csv.\n\nFor each series_id in test_images\nand for each image in series_id\nwe've to predict (5 conditions * 5 levels = 25) labels with normal, moderate and severe level.\n\nso submission format shouldn't be like this\nrow_id, series_id, conditions, normal, moderate and severe.",
    "3006505": "So we have to predict severity of each condition in each level. There sum up to 3 severity * 5 conditions * 5 levels = 75 that for 1 study (or 1 person/patient). Series is just different MRI planes of the same person. `row_id` is already included condition. Therefore, `row_id` is already enough. You can check out `sample_submission.csv`.",
    "3006696": "Thanks. But so far what I understood is that for each images in a study_id has share same labels associated with that study_id in train.csv. Is that correct? Can u please take a look on my code for x_train and y_train. \n\nAnd for submission.csv I still couldn't understand. Can u please further eleborate.\n\nHere's my code:\n**X_train = []\ny_train = []\n\nbase_train_dir = 'train_images'\n\ndef get_image_path(study_id, series_id, instance_number, base_dir):\n    return path + f\"{base_dir}/{study_id}/{series_id}/{instance_number}.dcm\"\n\ndef get_one_hot_labels(study_id):\n    \n    row = df_train[df_train['study_id'] == study_id]\n    \n    if not row.empty:\n        \n        labels = row.iloc[0, 1:].values\n        one_hot_labels = []\n        \n        for label in labels:\n            \n            if label == \"Normal/Mild\":\n                one_hot_labels.extend([1, 0, 0])\n            elif label == 'Moderate':\n                one_hot_labels.extend([0, 1, 0])\n            elif label == 'Severe':\n                one_hot_labels.extend([0, 0, 1])\n            else:\n                one_hot_labels.extend([0, 0, 0])\n        return one_hot_labels\n    else:\n        return None\n    \nfor _, row in df_train_label.iterrows():\n    study_id = row['study_id']\n    series_id = row['series_id']\n    instance_number = row['instance_number']\n    \n    image_path = get_image_path(study_id, series_id, instance_number, base_train_dir)\n    \n    ds = pydicom.dcmread(image_path)\n    image_data = ds.pixel_array\n    image_data = image_data.astype(np.float32) / np.max(image_data)\n    \n    image_data = cv2.resize(image_data, (224, 224))\n\n    X_train.append(image_data)\n    \n    labels = get_one_hot_labels(study_id)\n    if labels is not None:\n        y_train.append(labels)**",
    "3006742": ">for each images in a study_id has share same labels associated with that study_id in train.csv. Is that correct?\n\nYes, you are correct. \n\nHow you converting your labels is also depend on how you define you model. They seems fine if your CrossEntropy target is class probabilities. I recommend you study [this very well done notebook](https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline) especially in Dataset and train loop section. The author also demonstated dataloader which output class indices rather than class probabilities style. It might take time to understand but I think it is very time worthy going forward.",
    "3006763": "Thanks, I'll look into it. But Please help me if I need more help.",
    "3007066": "I've looked into it. It's a bit complicated for me. I've my own understanding... Please correct me.\n\nfor testing...\nlets assume that we've xyz study_id and 3 associated series_ids. Each series_id has different number of images. Lets say we've total of 20 images (from 3 series_id) and for each image I've to make prediction for 75 labels (in my case). So predictions array should've shape (20, 75). Correct?\n\nThen how to align this predictions array into form like submission.csv. Can u help me prepare submission.csv with this approach.\n\nCan u tell If I had chosen one of the correct solutions.",
    "3007236": "Most common way to do is stack images (1 image as 1 channel). For example, in the notebook, Author chose 10 for each series. So the final shape in 1 batch for individual study is [1, 30, 512, 512] (batch_size, channels, img_size, img_size). The output shape should be (batch_size, 75) in your case. 1 batch for 1 study. This is might be the most straight forward and easier solution. However, there is no correct solution. You need to try different approach for youself.",
    "3008582": "im kinda having some trouble trying to figure the submission format as well... \n\nfirst, how ima gonna associate each image to their respective condition+level if theres no such information on the csv file. they only refer to the series name. there are multiple images inside each series folder in the test, but each condition has 2 sides (left/right), except for spinal canal (sagittal t2/stir) and 5 levels! i know that certain image inside the subfolder is refered to a certain condition, but what side and what level? \n\nsecond, the fact there are multiple images on the test set also confuses me a bit... should i predict the probabilities for the severity scores( mild/moderate/severe) for all images then kinda average them?\n\nthank you for any clarifications! good luck yall."
  },
  "source": "meta"
}