{
  "id": 516802,
  "title": "ITK8191's Making dataset questions ",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516802",
  "author_name": "",
  "post_date": "2024-07-03T16:01:58.003905Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Here is a public notebook used to create a <a href=\"https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-making-dataset\" target=\"_blank\">dataset</a> This is publicly available. I had a few questions regarding the same. </p>\n<ul>\n<li>Why do we take subset of the data. (10 For Sagittal T1 and Sagittal T2) Why not convert everything and then take the needed number during training? </li>\n</ul>",
  "messages": [
    {
      "id": "2903126",
      "postDate": "07/03/2024 16:01:58",
      "content": "<p>Here is a public notebook used to create a <a href=\"https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-making-dataset\" target=\"_blank\">dataset</a> This is publicly available. I had a few questions regarding the same. </p>\n<ul>\n<li>Why do we take subset of the data. (10 For Sagittal T1 and Sagittal T2) Why not convert everything and then take the needed number during training? </li>\n</ul>",
      "rawMarkdown": "Here is a public notebook used to create a [dataset](https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-making-dataset) This is publicly available. I had a few questions regarding the same. \n- Why do we take subset of the data. (10 For Sagittal T1 and Sagittal T2) Why not convert everything and then take the needed number during training?",
      "votes": null
    },
    {
      "id": "2903259",
      "postDate": "07/03/2024 17:13:52",
      "content": "<p>You can maybe ping ITK8191, but my understanding would be that due to massive variation in the amount of data present study, by study and limitations in input model size, 30 slices (10 from each type) are uniformly sampled.<br>\nI did try modifying the dataset to increase sampling, but no improvement was visible on public leaderboard.</p>",
      "rawMarkdown": "You can maybe ping ITK8191, but my understanding would be that due to massive variation in the amount of data present study, by study and limitations in input model size, 30 slices (10 from each type) are uniformly sampled.\nI did try modifying the dataset to increase sampling, but no improvement was visible on public leaderboard.",
      "votes": null
    },
    {
      "id": "2903269",
      "postDate": "07/03/2024 17:22:11",
      "content": "<p>Oh okay. I understand that. Thank you for the reply</p>\n<p>I did ping him a couple of days back but did not get a response, so I thought someone here would let me know the reason :) </p>",
      "rawMarkdown": "Oh okay. I understand that. Thank you for the reply\n\nI did ping him a couple of days back but did not get a response, so I thought someone here would let me know the reason :)",
      "votes": null
    },
    {
      "id": "2914900",
      "postDate": "07/10/2024 08:30:05",
      "content": "<p>In the dataset from <a href=\"https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline\" target=\"_blank\">https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline</a>, I think it's better understood. What he wants to do is return an input composed of 30 layers, with ten of them being Sagittal T1, ten Sagittal T2, and ten Axial T2. The point is that he wants those ten images to be uniformly chosen, so for the Sagittal ones he can transform only the images that interest him. For Axial, he can't do this because he doesn't know their order, so he simply chooses 10 randomly. At least that's what I've understood from his proposal.</p>",
      "rawMarkdown": "In the dataset from https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline, I think it's better understood. What he wants to do is return an input composed of 30 layers, with ten of them being Sagittal T1, ten Sagittal T2, and ten Axial T2. The point is that he wants those ten images to be uniformly chosen, so for the Sagittal ones he can transform only the images that interest him. For Axial, he can't do this because he doesn't know their order, so he simply chooses 10 randomly. At least that's what I've understood from his proposal.",
      "votes": null
    },
    {
      "id": "2914998",
      "postDate": "07/10/2024 09:42:56",
      "content": "<p>That brings me back to the another question. What is this order ? </p>",
      "rawMarkdown": "That brings me back to the another question. What is this order ?",
      "votes": null
    },
    {
      "id": "2915899",
      "postDate": "07/10/2024 17:31:53",
      "content": "<p>When you do a CT scan, you get slices of a person's body in different planes (sagittal or axial in this competition). The numbers in the DICOM filenames indicate the order of the slices, not the parameter inside the DICOM file named something like slice_position. There is another discussion about the file order here: (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/518525)\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/518525)</a>.</p>",
      "rawMarkdown": "When you do a CT scan, you get slices of a person's body in different planes (sagittal or axial in this competition). The numbers in the DICOM filenames indicate the order of the slices, not the parameter inside the DICOM file named something like slice_position. There is another discussion about the file order here: (https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/518525).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2903259,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "07/03/2024 17:13:52",
      "content": "<p>You can maybe ping ITK8191, but my understanding would be that due to massive variation in the amount of data present study, by study and limitations in input model size, 30 slices (10 from each type) are uniformly sampled.<br>\nI did try modifying the dataset to increase sampling, but no improvement was visible on public leaderboard.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2903269,
          "author_name": "skaarface",
          "author_url": "",
          "post_date": "07/03/2024 17:22:11",
          "content": "<p>Oh okay. I understand that. Thank you for the reply</p>\n<p>I did ping him a couple of days back but did not get a response, so I thought someone here would let me know the reason :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2914900,
      "author_name": "pepe97",
      "author_url": "",
      "post_date": "07/10/2024 08:30:05",
      "content": "<p>In the dataset from <a href=\"https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline\" target=\"_blank\">https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline</a>, I think it's better understood. What he wants to do is return an input composed of 30 layers, with ten of them being Sagittal T1, ten Sagittal T2, and ten Axial T2. The point is that he wants those ten images to be uniformly chosen, so for the Sagittal ones he can transform only the images that interest him. For Axial, he can't do this because he doesn't know their order, so he simply chooses 10 randomly. At least that's what I've understood from his proposal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2914998,
          "author_name": "skaarface",
          "author_url": "",
          "post_date": "07/10/2024 09:42:56",
          "content": "<p>That brings me back to the another question. What is this order ? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2915899,
              "author_name": "pepe97",
              "author_url": "",
              "post_date": "07/10/2024 17:31:53",
              "content": "<p>When you do a CT scan, you get slices of a person's body in different planes (sagittal or axial in this competition). The numbers in the DICOM filenames indicate the order of the slices, not the parameter inside the DICOM file named something like slice_position. There is another discussion about the file order here: (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/518525)\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/518525)</a>.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2903126": "Here is a public notebook used to create a [dataset](https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-making-dataset) This is publicly available. I had a few questions regarding the same. \n- Why do we take subset of the data. (10 For Sagittal T1 and Sagittal T2) Why not convert everything and then take the needed number during training?",
    "2903259": "You can maybe ping ITK8191, but my understanding would be that due to massive variation in the amount of data present study, by study and limitations in input model size, 30 slices (10 from each type) are uniformly sampled.\nI did try modifying the dataset to increase sampling, but no improvement was visible on public leaderboard.",
    "2903269": "Oh okay. I understand that. Thank you for the reply\n\nI did ping him a couple of days back but did not get a response, so I thought someone here would let me know the reason :)",
    "2914900": "In the dataset from https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline, I think it's better understood. What he wants to do is return an input composed of 30 layers, with ten of them being Sagittal T1, ten Sagittal T2, and ten Axial T2. The point is that he wants those ten images to be uniformly chosen, so for the Sagittal ones he can transform only the images that interest him. For Axial, he can't do this because he doesn't know their order, so he simply chooses 10 randomly. At least that's what I've understood from his proposal.",
    "2914998": "That brings me back to the another question. What is this order ?",
    "2915899": "When you do a CT scan, you get slices of a person's body in different planes (sagittal or axial in this competition). The numbers in the DICOM filenames indicate the order of the slices, not the parameter inside the DICOM file named something like slice_position. There is another discussion about the file order here: (https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/518525)."
  },
  "source": "meta"
}