{
  "id": 517148,
  "title": "No of images used for inference",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/517148",
  "author_name": "",
  "post_date": "2024-07-05T03:24:46.911428800Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I see many notebooks and discussion on using samples of images rather than all the images for inference.<br>\nCan anyone explain why are we using samples and not whole image dataset and how are we selecting the images to train on ? </p>",
  "messages": [
    {
      "id": "2905539",
      "postDate": "07/05/2024 03:24:46",
      "content": "<p>I see many notebooks and discussion on using samples of images rather than all the images for inference.<br>\nCan anyone explain why are we using samples and not whole image dataset and how are we selecting the images to train on ? </p>",
      "rawMarkdown": "I see many notebooks and discussion on using samples of images rather than all the images for inference.\nCan anyone explain why are we using samples and not whole image dataset and how are we selecting the images to train on ?",
      "votes": null
    },
    {
      "id": "2906183",
      "postDate": "07/05/2024 12:18:27",
      "content": "<p>Each sub-folder contains a sequence of MRI images. For the training dataset, we have been provided an instance_number which corresponds to the image chosen by the radiologist for diagnosis. So for training, we could directly use these images. For inference, we again have sub-folders which contain a sequence of MRI images. We first need to identify which of these is the best one to provide to our model. One way of doing this would be to create a two-stage model, one for ROI detection and a second for classification. The ROI detector could help us choose the best image or images in the sequence. We could then feed this best image or top images to the classifier to make the predictions. In this case, while the ROI detector would use all the images in the sub-folders, the classifier will only use either one per sub-folder, or perhaps top-3 or top-5 per sub-folder.<br>\nThis is if we choose to train a 2D CNN for the classifier. I'm not sure, but maybe a 3D CNN could also be used which would then take in all the files in the MRI sequence to give predictions. But I guess this would be resource and time consuming for training as well as inference, so perhaps a two-stage model is easier to implement. </p>",
      "rawMarkdown": "Each sub-folder contains a sequence of MRI images. For the training dataset, we have been provided an instance_number which corresponds to the image chosen by the radiologist for diagnosis. So for training, we could directly use these images. For inference, we again have sub-folders which contain a sequence of MRI images. We first need to identify which of these is the best one to provide to our model. One way of doing this would be to create a two-stage model, one for ROI detection and a second for classification. The ROI detector could help us choose the best image or images in the sequence. We could then feed this best image or top images to the classifier to make the predictions. In this case, while the ROI detector would use all the images in the sub-folders, the classifier will only use either one per sub-folder, or perhaps top-3 or top-5 per sub-folder.\nThis is if we choose to train a 2D CNN for the classifier. I'm not sure, but maybe a 3D CNN could also be used which would then take in all the files in the MRI sequence to give predictions. But I guess this would be resource and time consuming for training as well as inference, so perhaps a two-stage model is easier to implement.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2906183,
      "author_name": "kaiserm",
      "author_url": "",
      "post_date": "07/05/2024 12:18:27",
      "content": "<p>Each sub-folder contains a sequence of MRI images. For the training dataset, we have been provided an instance_number which corresponds to the image chosen by the radiologist for diagnosis. So for training, we could directly use these images. For inference, we again have sub-folders which contain a sequence of MRI images. We first need to identify which of these is the best one to provide to our model. One way of doing this would be to create a two-stage model, one for ROI detection and a second for classification. The ROI detector could help us choose the best image or images in the sequence. We could then feed this best image or top images to the classifier to make the predictions. In this case, while the ROI detector would use all the images in the sub-folders, the classifier will only use either one per sub-folder, or perhaps top-3 or top-5 per sub-folder.<br>\nThis is if we choose to train a 2D CNN for the classifier. I'm not sure, but maybe a 3D CNN could also be used which would then take in all the files in the MRI sequence to give predictions. But I guess this would be resource and time consuming for training as well as inference, so perhaps a two-stage model is easier to implement. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2905539": "I see many notebooks and discussion on using samples of images rather than all the images for inference.\nCan anyone explain why are we using samples and not whole image dataset and how are we selecting the images to train on ?",
    "2906183": "Each sub-folder contains a sequence of MRI images. For the training dataset, we have been provided an instance_number which corresponds to the image chosen by the radiologist for diagnosis. So for training, we could directly use these images. For inference, we again have sub-folders which contain a sequence of MRI images. We first need to identify which of these is the best one to provide to our model. One way of doing this would be to create a two-stage model, one for ROI detection and a second for classification. The ROI detector could help us choose the best image or images in the sequence. We could then feed this best image or top images to the classifier to make the predictions. In this case, while the ROI detector would use all the images in the sub-folders, the classifier will only use either one per sub-folder, or perhaps top-3 or top-5 per sub-folder.\nThis is if we choose to train a 2D CNN for the classifier. I'm not sure, but maybe a 3D CNN could also be used which would then take in all the files in the MRI sequence to give predictions. But I guess this would be resource and time consuming for training as well as inference, so perhaps a two-stage model is easier to implement."
  },
  "source": "meta"
}