{
  "id": 274493,
  "title": "Questions about dataset and its preprocessing",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/274493",
  "author_name": "",
  "post_date": "2021-09-26T12:43:32.109977Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The main question is that channels of every patient have different number of images and because of that we can't construct a 3-d image for CNN. <br>\nThe second question is about techniques of preprocessing - is there any automatic way to throw useless images away from dataset and make the number of images in every channel equal at least for one patient?<br>\nMy idea was to compute mean gray value on image and delete the least gray images but it turned out to be wrong.</p>\n<p>I will be very grateful for any article or tip.</p>",
  "messages": [
    {
      "id": "1524344",
      "postDate": "09/26/2021 12:43:32",
      "content": "<p>The main question is that channels of every patient have different number of images and because of that we can't construct a 3-d image for CNN. <br>\nThe second question is about techniques of preprocessing - is there any automatic way to throw useless images away from dataset and make the number of images in every channel equal at least for one patient?<br>\nMy idea was to compute mean gray value on image and delete the least gray images but it turned out to be wrong.</p>\n<p>I will be very grateful for any article or tip.</p>",
      "rawMarkdown": "The main question is that channels of every patient have different number of images and because of that we can't construct a 3-d image for CNN. \nThe second question is about techniques of preprocessing - is there any automatic way to throw useless images away from dataset and make the number of images in every channel equal at least for one patient?\nMy idea was to compute mean gray value on image and delete the least gray images but it turned out to be wrong.\n\nI will be very grateful for any article or tip.",
      "votes": null
    },
    {
      "id": "1526163",
      "postDate": "09/27/2021 23:03:21",
      "content": "<p>On question 1: not only are there different numbers of images, they are also different sizes (FLAIR tends to be 512x512, T1w 256x256, etc.) and in different orientations ('Image Orientation (Patient)' is the relevant term in the DICOM files).  To access the 'Image Orientation (Patient)' data for a particular image, use:<br>\ndicom = pydicom.dcmread(image)<br>\norientation = dicom[('0020', '0037')].value<br>\nThis will return a list of numbers describing the orientation, although I haven't parsed out exactly what each component corresponds to physically.  I <em>have</em> determined that within a given mode for any patient all the images have the same orientation out to 4 decimal places, which is much less than 1 pixel for these images.  </p>\n<p>On question 2: Yes, you can throw away the (many) blank images fairly easily.  I found it useful to preprocess all of the patients, keeping only the images that have at least 1 non-zero pixel, and storing all this information in a JSON file, which I have linked: <a href=\"https://github.com/Astromax/RSNA-MICCAI-Brain-Tumor-Radiogenomic-Classification/blob/main/all_valid_image_names.json\" target=\"_blank\">https://github.com/Astromax/RSNA-MICCAI-Brain-Tumor-Radiogenomic-Classification/blob/main/all_valid_image_names.json</a> <br>\nI also sorted the images according to \"slice location\", so they would automatically be in order.  Note, there are 14 patients who DO NOT have slice location in their DICOM data, so they are excluded from my JSON file; these patients have IDs: '00109', '00157', '00170', '00186', '00353', '00367', '00414', '00561', '00563', '00564', '00565',<br>\n         '00756', '00834', '00839'</p>\n<p>Hope this is useful!</p>",
      "rawMarkdown": "On question 1: not only are there different numbers of images, they are also different sizes (FLAIR tends to be 512x512, T1w 256x256, etc.) and in different orientations ('Image Orientation (Patient)' is the relevant term in the DICOM files).  To access the 'Image Orientation (Patient)' data for a particular image, use:\ndicom = pydicom.dcmread(image)\norientation = dicom[('0020', '0037')].value\nThis will return a list of numbers describing the orientation, although I haven't parsed out exactly what each component corresponds to physically.  I *have* determined that within a given mode for any patient all the images have the same orientation out to 4 decimal places, which is much less than 1 pixel for these images.  \n\nOn question 2: Yes, you can throw away the (many) blank images fairly easily.  I found it useful to preprocess all of the patients, keeping only the images that have at least 1 non-zero pixel, and storing all this information in a JSON file, which I have linked: https://github.com/Astromax/RSNA-MICCAI-Brain-Tumor-Radiogenomic-Classification/blob/main/all_valid_image_names.json \nI also sorted the images according to \"slice location\", so they would automatically be in order.  Note, there are 14 patients who DO NOT have slice location in their DICOM data, so they are excluded from my JSON file; these patients have IDs: '00109', '00157', '00170', '00186', '00353', '00367', '00414', '00561', '00563', '00564', '00565',\n         '00756', '00834', '00839'\n\nHope this is useful!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1526163,
      "author_name": "maxbaugh",
      "author_url": "",
      "post_date": "09/27/2021 23:03:21",
      "content": "<p>On question 1: not only are there different numbers of images, they are also different sizes (FLAIR tends to be 512x512, T1w 256x256, etc.) and in different orientations ('Image Orientation (Patient)' is the relevant term in the DICOM files).  To access the 'Image Orientation (Patient)' data for a particular image, use:<br>\ndicom = pydicom.dcmread(image)<br>\norientation = dicom[('0020', '0037')].value<br>\nThis will return a list of numbers describing the orientation, although I haven't parsed out exactly what each component corresponds to physically.  I <em>have</em> determined that within a given mode for any patient all the images have the same orientation out to 4 decimal places, which is much less than 1 pixel for these images.  </p>\n<p>On question 2: Yes, you can throw away the (many) blank images fairly easily.  I found it useful to preprocess all of the patients, keeping only the images that have at least 1 non-zero pixel, and storing all this information in a JSON file, which I have linked: <a href=\"https://github.com/Astromax/RSNA-MICCAI-Brain-Tumor-Radiogenomic-Classification/blob/main/all_valid_image_names.json\" target=\"_blank\">https://github.com/Astromax/RSNA-MICCAI-Brain-Tumor-Radiogenomic-Classification/blob/main/all_valid_image_names.json</a> <br>\nI also sorted the images according to \"slice location\", so they would automatically be in order.  Note, there are 14 patients who DO NOT have slice location in their DICOM data, so they are excluded from my JSON file; these patients have IDs: '00109', '00157', '00170', '00186', '00353', '00367', '00414', '00561', '00563', '00564', '00565',<br>\n         '00756', '00834', '00839'</p>\n<p>Hope this is useful!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1524344": "The main question is that channels of every patient have different number of images and because of that we can't construct a 3-d image for CNN. \nThe second question is about techniques of preprocessing - is there any automatic way to throw useless images away from dataset and make the number of images in every channel equal at least for one patient?\nMy idea was to compute mean gray value on image and delete the least gray images but it turned out to be wrong.\n\nI will be very grateful for any article or tip.",
    "1526163": "On question 1: not only are there different numbers of images, they are also different sizes (FLAIR tends to be 512x512, T1w 256x256, etc.) and in different orientations ('Image Orientation (Patient)' is the relevant term in the DICOM files).  To access the 'Image Orientation (Patient)' data for a particular image, use:\ndicom = pydicom.dcmread(image)\norientation = dicom[('0020', '0037')].value\nThis will return a list of numbers describing the orientation, although I haven't parsed out exactly what each component corresponds to physically.  I *have* determined that within a given mode for any patient all the images have the same orientation out to 4 decimal places, which is much less than 1 pixel for these images.  \n\nOn question 2: Yes, you can throw away the (many) blank images fairly easily.  I found it useful to preprocess all of the patients, keeping only the images that have at least 1 non-zero pixel, and storing all this information in a JSON file, which I have linked: https://github.com/Astromax/RSNA-MICCAI-Brain-Tumor-Radiogenomic-Classification/blob/main/all_valid_image_names.json \nI also sorted the images according to \"slice location\", so they would automatically be in order.  Note, there are 14 patients who DO NOT have slice location in their DICOM data, so they are excluded from my JSON file; these patients have IDs: '00109', '00157', '00170', '00186', '00353', '00367', '00414', '00561', '00563', '00564', '00565',\n         '00756', '00834', '00839'\n\nHope this is useful!"
  },
  "source": "meta"
}