{
  "id": 253000,
  "title": "DICOM to PNG dataset (128 GB -> 5.2 GB) 🎨🔥",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/253000",
  "author_name": "Jonathan Besomi",
  "post_date": "2021-07-14T15:44:09.438000",
  "votes": 369,
  "comment_count": 103,
  "views": 0,
  "content": "<p>Preprocessing DICOM files is generally quite computational-intensive and slow, especially considering the large dataset size (~128 GB). For this reason, I've fired up a GCP VM and transformed the data from DICOM to PNG for you. The dataset file size has been reduced from 128 GB to 5.2 GB. </p>\n<p><strong>Notes</strong></p>\n<ul>\n<li>Images sizes have been kept as the original ones. </li>\n<li>To further reduce file size, all empty DICOM images files have not been included in the dataset, you can easily spot them by looking at the images sequences (<code>Image-X.png</code>). Note that, because of that, some entire folders (such as <code>train/00109/FLAIR</code>) are not included in this dataset (as they contain empty images only).</li>\n<li>The structure of the dataset is not changed. I.e all <em>.dcm</em> files are now <em>.png</em> files.</li>\n</ul>\n<p><strong>Dataset link</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/jonathanbesomi/rsna-miccai-png\" target=\"_blank\">RSNA-MICCAI-PNG-Dataset</a> [5.2 GB]</li>\n</ul>\n<p><strong>Script</strong></p>\n<p>For each image, run: (Adapted from <a href=\"https://www.kaggle.com/tanlikesmath\" target=\"_blank\">@tanlikesmath</a> <a href=\"https://www.kaggle.com/tanlikesmath/brain-tumor-radiogenomic-classification-eda\" target=\"_blank\">work</a>)</p>\n<pre><code>dicom = pydicom.read_file(path)\ndata = apply_voi_lut(dicom.pixel_array, dicom)\nif dicom.PhotometricInterpretation == \"MONOCHROME1\":\n    data = np.amax(data) - data\ndata = data - np.min(data)\ndata = data / np.max(data)\ndata = (data * 255).astype(np.uint8)\n</code></pre>\n<p><strong>More</strong></p>\n<p>If you need other formats, you want a version with resized images or anything else, just ask 👍</p>\n<p>Thank you for reading, happy kaggling to all!</p>",
  "messages": [
    {
      "id": 1388021,
      "postDate": "2021-07-14T15:44:09.440Z",
      "content": "<p>Preprocessing DICOM files is generally quite computational-intensive and slow, especially considering the large dataset size (~128 GB). For this reason, I've fired up a GCP VM and transformed the data from DICOM to PNG for you. The dataset file size has been reduced from 128 GB to 5.2 GB. </p>\n<p><strong>Notes</strong></p>\n<ul>\n<li>Images sizes have been kept as the original ones. </li>\n<li>To further reduce file size, all empty DICOM images files have not been included in the dataset, you can easily spot them by looking at the images sequences (<code>Image-X.png</code>). Note that, because of that, some entire folders (such as <code>train/00109/FLAIR</code>) are not included in this dataset (as they contain empty images only).</li>\n<li>The structure of the dataset is not changed. I.e all <em>.dcm</em> files are now <em>.png</em> files.</li>\n</ul>\n<p><strong>Dataset link</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/jonathanbesomi/rsna-miccai-png\" target=\"_blank\">RSNA-MICCAI-PNG-Dataset</a> [5.2 GB]</li>\n</ul>\n<p><strong>Script</strong></p>\n<p>For each image, run: (Adapted from <a href=\"https://www.kaggle.com/tanlikesmath\" target=\"_blank\">@tanlikesmath</a> <a href=\"https://www.kaggle.com/tanlikesmath/brain-tumor-radiogenomic-classification-eda\" target=\"_blank\">work</a>)</p>\n<pre><code>dicom = pydicom.read_file(path)\ndata = apply_voi_lut(dicom.pixel_array, dicom)\nif dicom.PhotometricInterpretation == \"MONOCHROME1\":\n    data = np.amax(data) - data\ndata = data - np.min(data)\ndata = data / np.max(data)\ndata = (data * 255).astype(np.uint8)\n</code></pre>\n<p><strong>More</strong></p>\n<p>If you need other formats, you want a version with resized images or anything else, just ask 👍</p>\n<p>Thank you for reading, happy kaggling to all!</p>",
      "rawMarkdown": "Preprocessing DICOM files is generally quite computational-intensive and slow, especially considering the large dataset size (~128 GB). For this reason, I've fired up a GCP VM and transformed the data from DICOM to PNG for you. The dataset file size has been reduced from 128 GB to 5.2 GB. \n\n**Notes**\n\n - Images sizes have been kept as the original ones. \n - To further reduce file size, all empty DICOM images files have not been included in the dataset, you can easily spot them by looking at the images sequences (`Image-X.png`). Note that, because of that, some entire folders (such as `train/00109/FLAIR`) are not included in this dataset (as they contain empty images only).\n - The structure of the dataset is not changed. I.e all _.dcm_ files are now _.png_ files.\n\n**Dataset link**\n\n - [RSNA-MICCAI-PNG-Dataset](https://www.kaggle.com/jonathanbesomi/rsna-miccai-png) [5.2 GB]\n\n**Script**\n\nFor each image, run: (Adapted from @tanlikesmath [work](https://www.kaggle.com/tanlikesmath/brain-tumor-radiogenomic-classification-eda))\n\n```\ndicom = pydicom.read_file(path)\ndata = apply_voi_lut(dicom.pixel_array, dicom)\nif dicom.PhotometricInterpretation == \"MONOCHROME1\":\n    data = np.amax(data) - data\ndata = data - np.min(data)\ndata = data / np.max(data)\ndata = (data * 255).astype(np.uint8)\n```\n\n**More**\n\nIf you need other formats, you want a version with resized images or anything else, just ask 👍\n\nThank you for reading, happy kaggling to all!",
      "votes": 369
    },
    {
      "id": 1395928,
      "postDate": "2021-07-21T16:14:41.843Z",
      "content": "<p>Greate work.</p>\n<p>Could you please elaborate on the decision to normalize an images the following way?</p>\n<pre><code>data = data - np.min(data)\ndata = data / np.max(data)\ndata = (data * 255).astype(np.uint8)\n</code></pre>",
      "rawMarkdown": "Greate work.\n\nCould you please elaborate on the decision to normalize an images the following way?\n```\ndata = data - np.min(data)\ndata = data / np.max(data)\ndata = (data * 255).astype(np.uint8)\n```",
      "votes": 3,
      "replies": [
        {
          "id": 1396936,
          "postDate": "2021-07-22T15:41:54.120Z",
          "content": "<p>saved images wont be a blank.</p>",
          "rawMarkdown": "saved images wont be a blank."
        }
      ]
    },
    {
      "id": 1389541,
      "postDate": "2021-07-15T19:52:48.717Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 3,
      "replies": [
        {
          "id": 1390192,
          "postDate": "2021-07-16T12:46:04.687Z",
          "content": "<p>Good luck with the comp. <a href=\"https://www.kaggle.com/cinthiakleiner\" target=\"_blank\">@cinthiakleiner</a> !!</p>",
          "rawMarkdown": "Good luck with the comp. @cinthiakleiner !!",
          "votes": -1
        }
      ]
    },
    {
      "id": 1392053,
      "postDate": "2021-07-18T10:33:41.790Z",
      "content": "<p>Thank you for sharing this..</p>",
      "rawMarkdown": "Thank you for sharing this..",
      "votes": 4,
      "replies": [
        {
          "id": 1393249,
          "postDate": "2021-07-19T13:48:42.510Z",
          "content": "<p>You are welcome, hope it helps! :)</p>",
          "rawMarkdown": "You are welcome, hope it helps! :)",
          "votes": -1
        }
      ]
    },
    {
      "id": 1471361,
      "postDate": "2021-08-14T06:50:43.293Z",
      "content": "<p>Why do you use this normalization formula?<br>\n<code>data = data - np.min(data)</code><br>\n  <code>data = data / np.max(data)</code><br>\nI checked the original Dicom file, some local minimal values are not equal to the global minimum value, so is the maximum value. The global minimal value is -32768.0 and the global maximum value is 65535.0 (I only checked the converted dicom files). But the minimum value of \"Image-118.dcm\" is -32754.13895939086, and the maximum value of \"Image-118.dcm\" is 16424.83312182741. Your formula threats the -32754 as 0 and 16424 as 255, I think it's wrong.<br>\nI am considering to use this formula:<br>\n<code>global_min, global_max = -32768.0, 65535.0</code><br>\n<code>data = data - global_min</code><br>\n    <code>data = data / global_max</code><br>\nThe converted images look similar, but they have different pixel values.</p>",
      "rawMarkdown": "Why do you use this normalization formula?\n`data = data - np.min(data)`\n  `data = data / np.max(data)`\nI checked the original Dicom file, some local minimal values are not equal to the global minimum value, so is the maximum value. The global minimal value is -32768.0 and the global maximum value is 65535.0 (I only checked the converted dicom files). But the minimum value of \"Image-118.dcm\" is -32754.13895939086, and the maximum value of \"Image-118.dcm\" is 16424.83312182741. Your formula threats the -32754 as 0 and 16424 as 255, I think it's wrong.\nI am considering to use this formula:\n`global_min, global_max = -32768.0, 65535.0`\n`data = data - global_min`\n    `data = data / global_max`\nThe converted images look similar, but they have different pixel values.",
      "votes": 1,
      "replies": [
        {
          "id": 1503278,
          "postDate": "2021-09-05T08:11:35.600Z",
          "content": "<p>I agree, your point is valid if dicom images have pixel range of [-32768, 65535] then it should map them to [0,255] not the one which is found in particular dcm file</p>",
          "rawMarkdown": "I agree, your point is valid if dicom images have pixel range of [-32768, 65535] then it should map them to [0,255] not the one which is found in particular dcm file"
        }
      ]
    },
    {
      "id": 1389296,
      "postDate": "2021-07-15T15:32:41.653Z",
      "content": "<p>Thanks a lot!<br>\nBut why using min-max normalize instead of normalize them using their full range (i.e. 0 to 2 ** bit_stored)??</p>",
      "rawMarkdown": "Thanks a lot!\nBut why using min-max normalize instead of normalize them using their full range (i.e. 0 to 2 ** bit_stored)??",
      "votes": 1,
      "replies": [
        {
          "id": 1390196,
          "postDate": "2021-07-16T12:48:50.603Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/terenceythsu\" target=\"_blank\">@terenceythsu</a>, not sure I got your question … why would you consider values between max and 2**<em>stored bit</em>?</p>",
          "rawMarkdown": "Hey @terenceythsu, not sure I got your question ... why would you consider values between max and 2**_stored bit_?",
          "votes": -2
        },
        {
          "id": 1391233,
          "postDate": "2021-07-17T12:11:16.277Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jonathanbesomi\" target=\"_blank\">@jonathanbesomi</a> , I think you got my point.<br>\nI just thought that the absolute value has its meaning, so maybe min-max normalize will destroy them.</p>",
          "rawMarkdown": "Hi @jonathanbesomi , I think you got my point.\nI just thought that the absolute value has its meaning, so maybe min-max normalize will destroy them.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1388314,
      "postDate": "2021-07-14T20:00:45.857Z",
      "content": "<p>Thanks for sharing! </p>",
      "rawMarkdown": "Thanks for sharing! ~~Just for those who don't know the size of the images. I just download them and check. The size of the png images is 512 x 512.~~",
      "votes": 1,
      "replies": [
        {
          "id": 1388339,
          "postDate": "2021-07-14T20:37:10.277Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/yus002\" target=\"_blank\">@yus002</a>, actually, the dataset <strong>contains images of different sizes</strong>, exactly as the original DICOM files. I basically transformed the data from \".dcm\" to \".png\", without changing the dimensions of the images (as this can be done quickly during training time and it wasn't the goal of this step). The \"expensive\" operations are loading DICOM files and extract the images, and that's why this dataset. Hope things are more clear now!  </p>",
          "rawMarkdown": "Hey @yus002, actually, the dataset **contains images of different sizes**, exactly as the original DICOM files. I basically transformed the data from \".dcm\" to \".png\", without changing the dimensions of the images (as this can be done quickly during training time and it wasn't the goal of this step). The \"expensive\" operations are loading DICOM files and extract the images, and that's why this dataset. Hope things are more clear now!  ",
          "votes": 3
        },
        {
          "id": 1388342,
          "postDate": "2021-07-14T20:43:39.397Z",
          "content": "<p>I see. Thanks!</p>",
          "rawMarkdown": "I see. Thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1388037,
      "postDate": "2021-07-14T15:56:11.977Z",
      "content": "<p>Wow, this is amazing <a href=\"https://www.kaggle.com/jonathanbesomi\" target=\"_blank\">@jonathanbesomi</a>! My computer has an SSD, so it's limited when it comes to space. Your solution will make this dataset way more accessible to people like me! Thanks a lot!</p>",
      "rawMarkdown": "Wow, this is amazing @jonathanbesomi! My computer has an SSD, so it's limited when it comes to space. Your solution will make this dataset way more accessible to people like me! Thanks a lot!",
      "votes": 1,
      "replies": [
        {
          "id": 1388046,
          "postDate": "2021-07-14T16:01:02.010Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/jonaslneri\" target=\"_blank\">@jonaslneri</a>, thank you for this comment, it means a lot to me to know that this work can be useful for others! Good luck with the competition! ;) 🏆</p>",
          "rawMarkdown": "Hey @jonaslneri, thank you for this comment, it means a lot to me to know that this work can be useful for others! Good luck with the competition! ;) 🏆",
          "votes": 1
        }
      ]
    },
    {
      "id": 1502064,
      "postDate": "2021-09-03T19:04:42.150Z",
      "content": "<p>Thanks a lot for sharing, this is very kind of you.</p>",
      "rawMarkdown": "Thanks a lot for sharing, this is very kind of you."
    },
    {
      "id": 1392161,
      "postDate": "2021-07-18T11:59:42.060Z",
      "content": "<p>Great work Thanks for sharing !</p>",
      "rawMarkdown": "Great work Thanks for sharing !",
      "votes": 2,
      "replies": [
        {
          "id": 1393248,
          "postDate": "2021-07-19T13:47:48.787Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/mohamedhanyyy\" target=\"_blank\">@mohamedhanyyy</a> ! :)</p>",
          "rawMarkdown": "Thank you @mohamedhanyyy ! :)",
          "votes": -1
        }
      ]
    },
    {
      "id": 1391397,
      "postDate": "2021-07-17T14:50:51.857Z",
      "content": "<p>Thanks for sharing ! Extremely useful</p>",
      "rawMarkdown": "Thanks for sharing ! Extremely useful",
      "votes": 2,
      "replies": [
        {
          "id": 1393253,
          "postDate": "2021-07-19T13:51:32.413Z",
          "content": "<p>Thank you Rithik! 🎉</p>",
          "rawMarkdown": "Thank you Rithik! 🎉"
        }
      ]
    },
    {
      "id": 1390308,
      "postDate": "2021-07-16T14:51:45.033Z",
      "content": "<p>Pretty amazing work, thank you for sharing with community.</p>",
      "rawMarkdown": "Pretty amazing work, thank you for sharing with community.",
      "votes": 2,
      "replies": [
        {
          "id": 1393259,
          "postDate": "2021-07-19T13:52:27.640Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/shangweichen\" target=\"_blank\">@shangweichen</a> !!</p>",
          "rawMarkdown": "Thank you @shangweichen !!"
        }
      ]
    },
    {
      "id": 1389534,
      "postDate": "2021-07-15T19:41:14.060Z",
      "content": "<p>Thank you very much for saving so much time and space.</p>",
      "rawMarkdown": "Thank you very much for saving so much time and space.",
      "votes": 2,
      "replies": [
        {
          "id": 1390195,
          "postDate": "2021-07-16T12:46:57.323Z",
          "content": "<p>Glad the dataset can be useful! All the best for the competition! 👍</p>",
          "rawMarkdown": "Glad the dataset can be useful! All the best for the competition! 👍"
        }
      ]
    },
    {
      "id": 1388851,
      "postDate": "2021-07-15T09:05:29.050Z",
      "content": "<p>Thanks for sharing …. magnificent reduce…</p>",
      "rawMarkdown": "Thanks for sharing .... magnificent reduce...",
      "votes": 2,
      "replies": [
        {
          "id": 1390204,
          "postDate": "2021-07-16T12:55:20.143Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/laxmankusuma\" target=\"_blank\">@laxmankusuma</a> !! All the best for this competition and the future ones!</p>",
          "rawMarkdown": "Thanks @laxmankusuma !! All the best for this competition and the future ones!"
        }
      ]
    },
    {
      "id": 1388074,
      "postDate": "2021-07-14T16:22:29.750Z",
      "content": "<p>Firstly, many thanks for your help. <br>\nI just had a beginner question, there are 4 subfolders with lots of images in each, for one data point or one patient, which one use for model training or which images to choose for model training</p>",
      "rawMarkdown": "Firstly, many thanks for your help. \nI just had a beginner question, there are 4 subfolders with lots of images in each, for one data point or one patient, which one use for model training or which images to choose for model training",
      "votes": 2,
      "replies": [
        {
          "id": 1388084,
          "postDate": "2021-07-14T16:31:05.223Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/tharun2001\" target=\"_blank\">@tharun2001</a>!</p>\n<p><strong>Short answer</strong>: use all images you can for training.</p>\n<p><strong>Long answer</strong>: this is a binary classification task at the \"case\" level. For each \"case\", there are 4 folders and each folder contains multiple scans (of different types, you can <a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/data\" target=\"_blank\">read more in the data tab here</a>). Note also that the same folder structure is also available for the test part. Therefore, it's suggested to model the problem so that we can exploit all data available, for instance, making 3D reconstructions of the scans.</p>\n<p>Hope it helps. Good luck!</p>",
          "rawMarkdown": "Hey @tharun2001!\n\n**Short answer**: use all images you can for training.\n\n**Long answer**: this is a binary classification task at the \"case\" level. For each \"case\", there are 4 folders and each folder contains multiple scans (of different types, you can [read more in the data tab here](https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/data)). Note also that the same folder structure is also available for the test part. Therefore, it's suggested to model the problem so that we can exploit all data available, for instance, making 3D reconstructions of the scans.\n\nHope it helps. Good luck!",
          "votes": 3
        }
      ]
    },
    {
      "id": 1547031,
      "postDate": "2021-10-16T16:27:27.277Z",
      "content": "<p>Thank you so much for doing this and sharing it with everyone! Really made my life easier!</p>",
      "rawMarkdown": "Thank you so much for doing this and sharing it with everyone! Really made my life easier!"
    },
    {
      "id": 1540781,
      "postDate": "2021-10-10T22:37:27.200Z",
      "content": "<p>\"data = (data * 255).astype(np.uint8)\" why using \"uint8\" dtype since decom image take value between 0 and 4096 ? dont we lose important information ? </p>",
      "rawMarkdown": "\"data = (data * 255).astype(np.uint8)\" why using \"uint8\" dtype since decom image take value between 0 and 4096 ? dont we lose important information ? "
    },
    {
      "id": 1524936,
      "postDate": "2021-09-27T02:26:15.343Z",
      "content": "<p>Thanks for the great data.</p>\n<p>I'm trying to insert a directory structure data into kaggle's dataset instead of a file, but I can't do it.<br>\nHow did you manage to insert the data in a directory structure?</p>",
      "rawMarkdown": "Thanks for the great data.\n\nI'm trying to insert a directory structure data into kaggle's dataset instead of a file, but I can't do it.\nHow did you manage to insert the data in a directory structure?\n\n\n"
    },
    {
      "id": 1507500,
      "postDate": "2021-09-09T09:01:15.880Z",
      "content": "<p>Thank you for sharing!  It's a very useful to use by Colaboratory.</p>",
      "rawMarkdown": "Thank you for sharing!  It's a very useful to use by Colaboratory."
    },
    {
      "id": 1506049,
      "postDate": "2021-09-07T19:14:08.987Z",
      "content": "<p>Thanks for the dataset! Do you have the full script for converting everything to PNG?</p>",
      "rawMarkdown": "Thanks for the dataset! Do you have the full script for converting everything to PNG?"
    },
    {
      "id": 1503215,
      "postDate": "2021-09-05T06:36:18.467Z",
      "content": "<p>Thanks for sharing, helps with the training time :)</p>",
      "rawMarkdown": "Thanks for sharing, helps with the training time :)"
    },
    {
      "id": 1499113,
      "postDate": "2021-09-01T12:55:26.383Z",
      "content": "<p>How did you filter out the empty DICOM images ? Could you share the script that you used to generate the dataset ?</p>",
      "rawMarkdown": "How did you filter out the empty DICOM images ? Could you share the script that you used to generate the dataset ?"
    },
    {
      "id": 1485498,
      "postDate": "2021-08-22T06:41:55.600Z",
      "content": "<p>How does this code actually work out deep down? I'm curious.</p>",
      "rawMarkdown": "How does this code actually work out deep down? I'm curious."
    },
    {
      "id": 1464832,
      "postDate": "2021-08-10T18:18:42.247Z",
      "content": "<p>How would I write the result <code>data</code> to a <code>.png</code> file?</p>",
      "rawMarkdown": "How would I write the result `data` to a `.png` file?",
      "replies": [
        {
          "id": 1464996,
          "postDate": "2021-08-10T19:46:12.593Z",
          "content": "<pre><code>def process_dicom(path, outpath):\n    dicom = pydicom.read_file(path)\n    data = apply_voi_lut(dicom.pixel_array, dicom)\n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n\n    height = len(data)\n    width = len(data[0])\n\n    pixels_out = []\n    for row in data:\n        pixels_out.extend(row)\n\n    image_out = Image.new('L', (width, height))\n    image_out.save(outpath)\n</code></pre>\n<p>Tried something like this, I think the <code>mode</code> code is wrong! I am just getting a black image as output.  'L' is an 8-bit greyscale mode, see <a href=\"https://pillow.readthedocs.io/en/stable/handbook/concepts.html#concept-modes\" target=\"_blank\">https://pillow.readthedocs.io/en/stable/handbook/concepts.html#concept-modes</a>.</p>",
          "rawMarkdown": "```\ndef process_dicom(path, outpath):\n    dicom = pydicom.read_file(path)\n    data = apply_voi_lut(dicom.pixel_array, dicom)\n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n    \n    height = len(data)\n    width = len(data[0])\n    \n    pixels_out = []\n    for row in data:\n        pixels_out.extend(row)\n    \n    image_out = Image.new('L', (width, height))\n    image_out.save(outpath)\n```\n\nTried something like this, I think the `mode` code is wrong! I am just getting a black image as output.  'L' is an 8-bit greyscale mode, see https://pillow.readthedocs.io/en/stable/handbook/concepts.html#concept-modes."
        },
        {
          "id": 1464998,
          "postDate": "2021-08-10T19:46:38.560Z",
          "content": "<p>was following this example: <a href=\"https://stackoverflow.com/questions/30943966/how-can-i-create-a-png-image-file-from-a-list-of-pixel-values-in-python\" target=\"_blank\">https://stackoverflow.com/questions/30943966/how-can-i-create-a-png-image-file-from-a-list-of-pixel-values-in-python</a> </p>",
          "rawMarkdown": "was following this example: https://stackoverflow.com/questions/30943966/how-can-i-create-a-png-image-file-from-a-list-of-pixel-values-in-python "
        },
        {
          "id": 1465015,
          "postDate": "2021-08-10T20:01:26.530Z",
          "content": "<p>Hah, forgot to do</p>\n<pre><code>    image_out.putdata(pixels_out)\n</code></pre>\n<p>After creating the PIL object! It works now!</p>",
          "rawMarkdown": "Hah, forgot to do\n```\n    image_out.putdata(pixels_out)\n\n```\nAfter creating the PIL object! It works now!"
        }
      ]
    },
    {
      "id": 1464692,
      "postDate": "2021-08-10T17:03:25.943Z",
      "content": "<p>Apparently RE: this link <a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/263612#1464675\" target=\"_blank\">https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/263612#1464675</a></p>\n<p>We can't actually use the PNG dataset when submitting we have to convert from DICOM to PNG ourselves.</p>\n<p>Regardless this is very helpful, thank you!</p>",
      "rawMarkdown": "Apparently RE: this link https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/263612#1464675\n\nWe can't actually use the PNG dataset when submitting we have to convert from DICOM to PNG ourselves.\n\nRegardless this is very helpful, thank you!"
    },
    {
      "id": 1461562,
      "postDate": "2021-08-09T13:09:33.177Z",
      "content": "<p>Cool! <br>\nBut it is necessary to understand, that DICOM contains not only an image, but also different additional data. <a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/252942\" target=\"_blank\">Here</a> Peter showed how to extract it.</p>",
      "rawMarkdown": "Cool! \nBut it is necessary to understand, that DICOM contains not only an image, but also different additional data. [Here](https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/252942) Peter showed how to extract it."
    },
    {
      "id": 1460162,
      "postDate": "2021-08-08T17:56:04.167Z",
      "content": "<p>Awesome!!, Thank you so much for sharing this! :)</p>",
      "rawMarkdown": "Awesome!!, Thank you so much for sharing this! :)"
    },
    {
      "id": 1407313,
      "postDate": "2021-08-01T18:08:22.900Z",
      "content": "<p>Cool. Thats really helpful</p>",
      "rawMarkdown": "Cool. Thats really helpful"
    },
    {
      "id": 1403601,
      "postDate": "2021-07-29T09:11:54.803Z",
      "content": "<p>Hey,</p>\n<p>Thank you for the effort but I am curious, doesn't this conversion compromise the Meta-Data? Also, our inputs would be DICOM images, right?</p>",
      "rawMarkdown": "Hey,\n\nThank you for the effort but I am curious, doesn't this conversion compromise the Meta-Data? Also, our inputs would be DICOM images, right?"
    },
    {
      "id": 1403028,
      "postDate": "2021-07-28T18:12:58.263Z",
      "content": "<p>Hii, thank you for sharing your work with us. Can you also tell me how do I import the reduced dataset in google colab notebooks?</p>",
      "rawMarkdown": "Hii, thank you for sharing your work with us. Can you also tell me how do I import the reduced dataset in google colab notebooks?"
    },
    {
      "id": 1402521,
      "postDate": "2021-07-28T08:57:15.930Z",
      "content": "<p>Excellent!</p>",
      "rawMarkdown": "Excellent!"
    },
    {
      "id": 1401825,
      "postDate": "2021-07-27T15:32:43.177Z",
      "content": "<p>Thank you for sharing great &amp; helpful information!</p>",
      "rawMarkdown": "Thank you for sharing great & helpful information!"
    },
    {
      "id": 1398191,
      "postDate": "2021-07-23T20:15:10.140Z",
      "content": "<p>Thank you to share your knowledge! Great work!! </p>",
      "rawMarkdown": "Thank you to share your knowledge! Great work!! "
    },
    {
      "id": 1395919,
      "postDate": "2021-07-21T16:06:00.257Z",
      "content": "<p>Is there any link to the data where just format changed from dicom to png/jpeg, keeping other things exactly same, like not removing those empty image files etc ?<br>\nIt would be a great help, thanks.</p>",
      "rawMarkdown": "Is there any link to the data where just format changed from dicom to png/jpeg, keeping other things exactly same, like not removing those empty image files etc ?\nIt would be a great help, thanks."
    },
    {
      "id": 1395298,
      "postDate": "2021-07-21T05:46:41.627Z",
      "content": "<p>Thank for your help!!!<br>\nCan you share more sizes? More sizes will be very interesting!</p>",
      "rawMarkdown": "Thank for your help!!!\nCan you share more sizes? More sizes will be very interesting!"
    },
    {
      "id": 1395281,
      "postDate": "2021-07-21T05:23:05.650Z",
      "content": "<p>Thanks for sharing I was looking for this and this will be really helpful for me.</p>",
      "rawMarkdown": "Thanks for sharing I was looking for this and this will be really helpful for me."
    },
    {
      "id": 1394957,
      "postDate": "2021-07-20T18:49:56.077Z",
      "content": "<p>Great work.</p>",
      "rawMarkdown": "Great work."
    },
    {
      "id": 1394510,
      "postDate": "2021-07-20T12:21:57.700Z",
      "content": "<p>Sincerely Thanks for sharing this it will really help us in training our model</p>",
      "rawMarkdown": "Sincerely Thanks for sharing this it will really help us in training our model"
    },
    {
      "id": 1391993,
      "postDate": "2021-07-18T09:21:30.020Z",
      "content": "<p>Glad to see something I was looking for quite a while</p>",
      "rawMarkdown": "Glad to see something I was looking for quite a while"
    },
    {
      "id": 1391000,
      "postDate": "2021-07-17T08:22:26.017Z",
      "content": "<p>That is brilliant. Thank you for time and energy-saving tip, John!</p>",
      "rawMarkdown": "That is brilliant. Thank you for time and energy-saving tip, John!",
      "replies": [
        {
          "id": 1393258,
          "postDate": "2021-07-19T13:52:11.523Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/samerrkhann\" target=\"_blank\">@samerrkhann</a>! Good luck with this competition! 💥</p>",
          "rawMarkdown": "Thank you @samerrkhann! Good luck with this competition! 💥"
        }
      ]
    },
    {
      "id": 1390999,
      "postDate": "2021-07-17T08:22:20.667Z",
      "content": "<p>That is brilliant. Thank you for time and energy-saving tip, John!</p>",
      "rawMarkdown": "That is brilliant. Thank you for time and energy-saving tip, John!"
    },
    {
      "id": 1389165,
      "postDate": "2021-07-15T13:44:39.600Z",
      "content": "<p>Thanks for the transformation! any possibility of having the data in .nii.gz format? I am getting 3 different volumes for each image when applying dcm2nii</p>",
      "rawMarkdown": "Thanks for the transformation! any possibility of having the data in .nii.gz format? I am getting 3 different volumes for each image when applying dcm2nii",
      "replies": [
        {
          "id": 1390199,
          "postDate": "2021-07-16T12:51:22.003Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/ranafago\" target=\"_blank\">@ranafago</a>, what do you mean by this sentence: <em>I am getting 3 different volumes for each image when applying dcm2nii</em>? <br>\nI can try to look into what you are asking if there is a necessity .. why would you need data in the <em>.nii.gz</em> format? </p>",
          "rawMarkdown": "Hey @ranafago, what do you mean by this sentence: _I am getting 3 different volumes for each image when applying dcm2nii_? \nI can try to look into what you are asking if there is a necessity .. why would you need data in the _.nii.gz_ format? "
        },
        {
          "id": 1392583,
          "postDate": "2021-07-18T20:47:35.177Z",
          "content": "<p>I was working towards a model using handcrafted features since I expected a similar number of cases as last year's survival prediction task, and I just worked always with the data in .nii.gz. I understand the benefits of working with the data without the preprocessing, but I just want to make a small test to check if that approach is still viable against deep learning approaches. Nevertheless, the data for the segmentation task is already in that format, Thanks for your answer! </p>\n<p>I am getting this output from dcm2nii I understand that the same image is reconstructed cropped and reoriented too. FLAIR image of case 00000:</p>\n<p>Chris Rorden's dcm2nii :: 4AUGUST2014 (Debian) 64bit BSD License<br>\nreading preferences file /home/rafa/.dcm2nii/dcm2nii.ini<br>\nData will be exported to FLAIR<br>\nValidating 400 potential DICOM images.<br>\nFound 400 DICOM images.<br>\nConverting 400/400  volumes: 1<br>\nImage-1.dcm-&gt;18991230_000000s004a1000.nii<br>\nGZip…18991230_000000s004a1000.nii.gz<br>\nReorienting as FLAIR/o18991230_000000s004a1000.nii.gz<br>\nGZip…o18991230_000000s004a1000.nii.gz<br>\nCropping NIfTI/Analyze image FLAIR/o18991230_000000s004a1000.nii.gz<br>\nGZip…co18991230_000000s004a1000.nii.gz</p>",
          "rawMarkdown": "I was working towards a model using handcrafted features since I expected a similar number of cases as last year's survival prediction task, and I just worked always with the data in .nii.gz. I understand the benefits of working with the data without the preprocessing, but I just want to make a small test to check if that approach is still viable against deep learning approaches. Nevertheless, the data for the segmentation task is already in that format, Thanks for your answer! \n\nI am getting this output from dcm2nii I understand that the same image is reconstructed cropped and reoriented too. FLAIR image of case 00000:\n\nChris Rorden's dcm2nii :: 4AUGUST2014 (Debian) 64bit BSD License\nreading preferences file /home/rafa/.dcm2nii/dcm2nii.ini\nData will be exported to FLAIR\nValidating 400 potential DICOM images.\nFound 400 DICOM images.\nConverting 400/400  volumes: 1\nImage-1.dcm->18991230_000000s004a1000.nii\nGZip...18991230_000000s004a1000.nii.gz\nReorienting as FLAIR/o18991230_000000s004a1000.nii.gz\nGZip...o18991230_000000s004a1000.nii.gz\nCropping NIfTI/Analyze image FLAIR/o18991230_000000s004a1000.nii.gz\nGZip...co18991230_000000s004a1000.nii.gz\n\n "
        },
        {
          "id": 1402406,
          "postDate": "2021-07-28T07:04:35.300Z",
          "content": "<p>Dear Jonathan</p>\n<p>Kindly provide us a (python) code to convert the dicom files to nii.gz format. This will be quite helpful as the codes we have developed so far, with the BraTS dataset, are all written for nii.gz format. Your help is much appreciated.</p>\n<p>Kind Regards</p>",
          "rawMarkdown": "Dear Jonathan\n\nKindly provide us a (python) code to convert the dicom files to nii.gz format. This will be quite helpful as the codes we have developed so far, with the BraTS dataset, are all written for nii.gz format. Your help is much appreciated.\n\nKind Regards"
        },
        {
          "id": 1402463,
          "postDate": "2021-07-28T08:15:14.090Z",
          "content": "<p>Hey HMD! </p>\n<p>I did something similar here, check it out maybe it helps </p>\n<p><a href=\"https://www.kaggle.com/ranafago/preprocessing-dcm-to-nifti\" target=\"_blank\">https://www.kaggle.com/ranafago/preprocessing-dcm-to-nifti</a></p>",
          "rawMarkdown": "Hey HMD! \n\nI did something similar here, check it out maybe it helps \n\nhttps://www.kaggle.com/ranafago/preprocessing-dcm-to-nifti"
        },
        {
          "id": 1407359,
          "postDate": "2021-08-01T19:08:29.263Z",
          "content": "<p>Hi Ranafago,</p>\n<p>Thank you for sharing your code.</p>\n<p>I was wondering if you had to disable any flags while using the function convert_directory from dicom2nifti. For example I had to disable validate_orthogonal to get T2w.nii.gz for two of the subjects.</p>",
          "rawMarkdown": "Hi Ranafago,\n\nThank you for sharing your code.\n\nI was wondering if you had to disable any flags while using the function convert_directory from dicom2nifti. For example I had to disable validate_orthogonal to get T2w.nii.gz for two of the subjects.\n\n\n"
        }
      ]
    },
    {
      "id": 1389162,
      "postDate": "2021-07-15T13:42:21.547Z",
      "content": "<p>You saved my internet. And my bill. Thanks!</p>",
      "rawMarkdown": "You saved my internet. And my bill. Thanks!",
      "replies": [
        {
          "id": 1390202,
          "postDate": "2021-07-16T12:54:44.793Z",
          "content": "<p>Yeah, I had to use a virtual machine to download and unzip the data, way too much data for my current machine too! Happy it helped and have fun in this competition. Will you attempt to reach a second gold medal?</p>",
          "rawMarkdown": "Yeah, I had to use a virtual machine to download and unzip the data, way too much data for my current machine too! Happy it helped and have fun in this competition. Will you attempt to reach a second gold medal?",
          "votes": 1
        },
        {
          "id": 1391999,
          "postDate": "2021-07-18T09:28:22.420Z",
          "content": "<p>I am exploring the datasets and have my fun with personal projects. Such, compressed data is super helpful and experimental on my local machine 👍</p>",
          "rawMarkdown": "I am exploring the datasets and have my fun with personal projects. Such, compressed data is super helpful and experimental on my local machine 👍"
        }
      ]
    },
    {
      "id": 1388350,
      "postDate": "2021-07-14T21:06:12.237Z",
      "content": "<p>You have a place reserved in Heaven.</p>",
      "rawMarkdown": "You have a place reserved in Heaven.",
      "replies": [
        {
          "id": 1390210,
          "postDate": "2021-07-16T13:00:52.087Z",
          "content": "<p>That's a bit too much <a href=\"https://www.kaggle.com/josepc\" target=\"_blank\">@josepc</a> :D </p>",
          "rawMarkdown": "That's a bit too much @josepc :D "
        }
      ]
    },
    {
      "id": 1485282,
      "postDate": "2021-08-22T00:15:39.883Z",
      "content": "<p>Request to add train label file too😄. Thanks for your efforts you made to convert the file in .png file. Because of your efforts, it's now practical possible for me to work with. </p>",
      "rawMarkdown": "Request to add train label file too😄. Thanks for your efforts you made to convert the file in .png file. Because of your efforts, it's now practical possible for me to work with. ",
      "votes": 4,
      "isDeleted": true
    },
    {
      "id": 1400510,
      "postDate": "2021-07-26T11:02:11.120Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1395003,
      "postDate": "2021-07-20T19:46:17.440Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1393109,
      "postDate": "2021-07-19T11:40:30.547Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1393244,
          "postDate": "2021-07-19T13:44:42.103Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/aryamansharma47\" target=\"_blank\">@aryamansharma47</a>, thank you for your comment, this is not a mistake and it has been already explained <a href=\"https://www.kaggle.com/jonathanbesomi/rsna-miccai-png/discussion/253022\" target=\"_blank\">here</a>.</p>\n<p>Basically: as you can read from notes, \"<strong>to further reduce file size, all empty DICOM images files have not been included in the dataset</strong>\". All <em>FLAIR</em> folders for the folders you mentioned (train/00109, train/00123, etc) contain only empty images and therefore have been skipped.</p>\n<p>Hope it helps, good luck with the competition!</p>",
          "rawMarkdown": "Hi @aryamansharma47, thank you for your comment, this is not a mistake and it has been already explained [here](https://www.kaggle.com/jonathanbesomi/rsna-miccai-png/discussion/253022).\n\nBasically: as you can read from notes, \"**to further reduce file size, all empty DICOM images files have not been included in the dataset**\". All _FLAIR_ folders for the folders you mentioned (train/00109, train/00123, etc) contain only empty images and therefore have been skipped.\n\nHope it helps, good luck with the competition!",
          "votes": 1
        },
        {
          "id": 1393823,
          "postDate": "2021-07-20T00:09:33.667Z",
          "content": "<p>my bad haha, thanks </p>",
          "rawMarkdown": "my bad haha, thanks "
        }
      ]
    },
    {
      "id": 1392050,
      "postDate": "2021-07-18T10:28:23.657Z",
      "content": "<p>Thanks for the sharing. As the target is radiomics, is precision lost (from 16 bits to 8bits) can damage subtitles textural patterns ? :s </p>",
      "rawMarkdown": "Thanks for the sharing. As the target is radiomics, is precision lost (from 16 bits to 8bits) can damage subtitles textural patterns ? :s ",
      "isDeleted": true,
      "replies": [
        {
          "id": 1393252,
          "postDate": "2021-07-19T13:51:00.733Z",
          "content": "<p>Hi Stephan, yes, as already discussed below, as we go from 128 GB to ~5GB it's expected to have some information loss. It's the tradeoff between preprocessing time/file size and information loss.   </p>",
          "rawMarkdown": "Hi Stephan, yes, as already discussed below, as we go from 128 GB to ~5GB it's expected to have some information loss. It's the tradeoff between preprocessing time/file size and information loss.   "
        }
      ]
    },
    {
      "id": 1389521,
      "postDate": "2021-07-15T19:22:43.463Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1388414,
      "postDate": "2021-07-14T23:28:34.737Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1388747,
          "postDate": "2021-07-15T07:36:08.020Z",
          "content": "<p>Not exactly, what he does is convert the values to the 0-255 range by minmax scaling</p>",
          "rawMarkdown": "Not exactly, what he does is convert the values to the 0-255 range by minmax scaling"
        },
        {
          "id": 1389579,
          "postDate": "2021-07-15T20:51:14.420Z",
          "content": "<p>In that case, since the images are converted to int8 arrays, we are losing some data precision due to the dynamic range compression and rounding, right?</p>",
          "rawMarkdown": "In that case, since the images are converted to int8 arrays, we are losing some data precision due to the dynamic range compression and rounding, right?"
        },
        {
          "id": 1390205,
          "postDate": "2021-07-16T12:59:31.167Z",
          "content": "<blockquote>\n  <p>In that case, since the images are converted to int8 arrays, we are losing some data precision due to the dynamic range compression and rounding, right?</p>\n</blockquote>\n<p>It's a tradeoff. Surely we lost some data, as we went from 128 GB to ~6GB but the advantage is that the dataset is much more malleable right now. I believe still, we can build quite strong models on this dataset too, especially in the initial phase for faster iterations.</p>",
          "rawMarkdown": "> In that case, since the images are converted to int8 arrays, we are losing some data precision due to the dynamic range compression and rounding, right?\n\nIt's a tradeoff. Surely we lost some data, as we went from 128 GB to ~6GB but the advantage is that the dataset is much more malleable right now. I believe still, we can build quite strong models on this dataset too, especially in the initial phase for faster iterations.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1538355,
      "postDate": "2021-10-08T10:34:06.897Z",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!"
    },
    {
      "id": 1535553,
      "postDate": "2021-10-06T00:37:46.337Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 1508350,
      "postDate": "2021-09-10T06:59:43.050Z",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!"
    },
    {
      "id": 1507707,
      "postDate": "2021-09-09T13:03:12.023Z",
      "content": "<p>Thank you for sharing :) </p>",
      "rawMarkdown": "Thank you for sharing :) "
    },
    {
      "id": 1490981,
      "postDate": "2021-08-26T04:15:24.700Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!"
    },
    {
      "id": 1485310,
      "postDate": "2021-08-22T01:42:25.983Z",
      "content": "<p>Thank you  for sharing!!</p>",
      "rawMarkdown": "Thank you  for sharing!!"
    },
    {
      "id": 1477503,
      "postDate": "2021-08-17T13:59:11.910Z",
      "content": "<p>Awesome work! Thank you for sharing!</p>",
      "rawMarkdown": "Awesome work! Thank you for sharing!"
    },
    {
      "id": 1473915,
      "postDate": "2021-08-15T19:17:01.260Z",
      "content": "<p>thanks for sharing great work</p>",
      "rawMarkdown": "thanks for sharing great work"
    },
    {
      "id": 1473423,
      "postDate": "2021-08-15T15:05:44.153Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!"
    },
    {
      "id": 1473137,
      "postDate": "2021-08-15T11:41:05.720Z",
      "content": "<p>Thanks for sharing! very helpful!</p>",
      "rawMarkdown": "Thanks for sharing! very helpful!"
    },
    {
      "id": 1457486,
      "postDate": "2021-08-07T12:40:18.997Z",
      "content": "<p>thank you so much</p>",
      "rawMarkdown": "thank you so much\n"
    },
    {
      "id": 1443214,
      "postDate": "2021-08-04T02:12:53.270Z",
      "content": "<p>Thank you you saved my disk…. :)</p>",
      "rawMarkdown": "Thank you you saved my disk.... :)"
    },
    {
      "id": 1400843,
      "postDate": "2021-07-26T16:25:57.900Z",
      "content": "<p>Thanks a lot !</p>",
      "rawMarkdown": "Thanks a lot !"
    },
    {
      "id": 1400319,
      "postDate": "2021-07-26T07:54:06.937Z",
      "content": "<p>Thank you for sharing :)</p>",
      "rawMarkdown": "Thank you for sharing :)"
    },
    {
      "id": 1400037,
      "postDate": "2021-07-26T00:54:52.077Z",
      "content": "<p>Thank you Jonathan :D</p>",
      "rawMarkdown": "Thank you Jonathan :D"
    },
    {
      "id": 1397491,
      "postDate": "2021-07-23T08:41:11.983Z",
      "content": "<p>Great and Thanks for sharing this .</p>",
      "rawMarkdown": "Great and Thanks for sharing this .\n"
    },
    {
      "id": 1396288,
      "postDate": "2021-07-22T01:43:03.203Z",
      "content": "<p>Great! Thanks!</p>",
      "rawMarkdown": "Great! Thanks!"
    },
    {
      "id": 1395033,
      "postDate": "2021-07-20T20:27:02.627Z",
      "content": "<p>Thank you so much for sharing.</p>",
      "rawMarkdown": "Thank you so much for sharing."
    },
    {
      "id": 1393896,
      "postDate": "2021-07-20T02:25:59.653Z",
      "content": "<p>Thank you for such a nice work :)</p>",
      "rawMarkdown": "Thank you for such a nice work :)"
    },
    {
      "id": 1391078,
      "postDate": "2021-07-17T09:12:21.437Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing."
    }
  ],
  "comments": [
    {
      "id": 1395928,
      "author_name": "Roman",
      "author_url": "",
      "post_date": "2021-07-21T16:14:41.843000",
      "content": "<p>Greate work.</p>\n<p>Could you please elaborate on the decision to normalize an images the following way?</p>\n<pre><code>data = data - np.min(data)\ndata = data / np.max(data)\ndata = (data * 255).astype(np.uint8)\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 1396936,
          "author_name": "Neo",
          "author_url": "",
          "post_date": "2021-07-22T15:41:54.120000",
          "content": "<p>saved images wont be a blank.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1389541,
      "author_name": "Cinthia Cristina Calchi Kleiner",
      "author_url": "",
      "post_date": "2021-07-15T19:52:48.717000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1390192,
          "author_name": "Jonathan Besomi",
          "author_url": "",
          "post_date": "2021-07-16T12:46:04.687000",
          "content": "<p>Good luck with the comp. <a href=\"https://www.kaggle.com/cinthiakleiner\" target=\"_blank\">@cinthiakleiner</a> !!</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 1392053,
      "author_name": "Rahul Bhimani",
      "author_url": "",
      "post_date": "2021-07-18T10:33:41.790000",
      "content": "<p>Thank you for sharing this..</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1393249,
          "author_name": "Jonathan Besomi",
          "author_url": "",
          "post_date": "2021-07-19T13:48:42.510000",
          "content": "<p>You are welcome, hope it helps! :)</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 1471361,
      "author_name": "Shanshan Yu",
      "author_url": "",
      "post_date": "2021-08-14T06:50:43.293000",
      "content": "<p>Why do you use this normalization formula?<br>\n<code>data = data - np.min(data)</code><br>\n  <code>data = data / np.max(data)</code><br>\nI checked the original Dicom file, some local minimal values are not equal to the global minimum value, so is the maximum value. The global minimal value is -32768.0 and the global maximum value is 65535.0 (I only checked the converted dicom files). But the minimum value of \"Image-118.dcm\" is -32754.13895939086, and the maximum value of \"Image-118.dcm\" is 16424.83312182741. Your formula threats the -32754 as 0 and 16424 as 255, I think it's wrong.<br>\nI am considering to use this formula:<br>\n<code>global_min, global_max = -32768.0, 65535.0</code><br>\n<code>data = data - global_min</code><br>\n    <code>data = data / global_max</code><br>\nThe converted images look similar, but they have different pixel values.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1503278,
          "author_name": "Shrijeet16",
          "author_url": "",
          "post_date": "2021-09-05T08:11:35.600000",
          "content": "<p>I agree, your point is valid if dicom images have pixel range of [-32768, 65535] then it should map them to [0,255] not the one which is found in particular dcm file</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1389296,
      "author_name": "lazyterence",
      "author_url": "",
      "post_date": "2021-07-15T15:32:41.653000",
      "content": "<p>Thanks a lot!<br>\nBut why using min-max normalize instead of normalize them using their full range (i.e. 0 to 2 ** bit_stored)??</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1390196,
          "author_name": "Jonathan Besomi",
          "author_url": "",
          "post_date": "2021-07-16T12:48:50.603000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/terenceythsu\" target=\"_blank\">@terenceythsu</a>, not sure I got your question … why would you consider values between max and 2**<em>stored bit</em>?</p>",
          "votes": -2,
          "replies": []
        },
        {
          "id": 1391233,
          "author_name": "lazyterence",
          "author_url": "",
          "post_date": "2021-07-17T12:11:16.277000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jonathanbesomi\" target=\"_blank\">@jonathanbesomi</a> , I think you got my point.<br>\nI just thought that the absolute value has its meaning, so maybe min-max normalize will destroy them.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1388314,
      "author_name": "Yue Sun",
      "author_url": "",
      "post_date": "2021-07-14T20:00:45.857000",
      "content": "<p>Thanks for sharing! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1388339,
          "author_name": "Jonathan Besomi",
          "author_url": "",
          "post_date": "2021-07-14T20:37:10.277000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/yus002\" target=\"_blank\">@yus002</a>, actually, the dataset <strong>contains images of different sizes</strong>, exactly as the original DICOM files. I basically transformed the data from \".dcm\" to \".png\", without changing the dimensions of the images (as this can be done quickly during training time and it wasn't the goal of this step). The \"expensive\" operations are loading DICOM files and extract the images, and that's why this dataset. Hope things are more clear now!  </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1388342,
          "author_name": "Yue Sun",
          "author_url": "",
          "post_date": "2021-07-14T20:43:39.397000",
          "content": "<p>I see. Thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1388037,
      "author_name": "Jonas Neri",
      "author_url": "",
      "post_date": "2021-07-14T15:56:11.977000",
      "content": "<p>Wow, this is amazing <a href=\"https://www.kaggle.com/jonathanbesomi\" target=\"_blank\">@jonathanbesomi</a>! My computer has an SSD, so it's limited when it comes to space. Your solution will make this dataset way more accessible to people like me! Thanks a lot!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1388046,
          "author_name": "Jonathan Besomi",
          "author_url": "",
          "post_date": "2021-07-14T16:01:02.010000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/jonaslneri\" target=\"_blank\">@jonaslneri</a>, thank you for this comment, it means a lot to me to know that this work can be useful for others! Good luck with the competition! ;) 🏆</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1502064,
      "author_name": "Mahmoud Limam",
      "author_url": "",
      "post_date": "2021-09-03T19:04:42.150000",
      "content": "<p>Thanks a lot for sharing, this is very kind of you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1392161,
      "author_name": "Mohamed Hany",
      "author_url": "",
      "post_date": "2021-07-18T11:59:42.060000",
      "content": "<p>Great work Thanks for sharing !</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1393248,
          "author_name": "Jonathan Besomi",
          "author_url": "",
          "post_date": "2021-07-19T13:47:48.787000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/mohamedhanyyy\" target=\"_blank\">@mohamedhanyyy</a> ! :)</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 1391397,
      "author_name": "Rithik Kotha",
      "author_url": "",
      "post_date": "2021-07-17T14:50:51.857000",
      "content": "<p>Thanks for sharing ! Extremely useful</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1393253,
          "author_name": "Jonathan Besomi",
          "author_url": "",
          "post_date": "2021-07-19T13:51:32.413000",
          "content": "<p>Thank you Rithik! 🎉</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1390308,
      "author_name": "README",
      "author_url": "",
      "post_date": "2021-07-16T14:51:45.033000",
      "content": "<p>Pretty amazing work, thank you for sharing with community.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1393259,
          "author_name": "Jonathan Besomi",
          "author_url": "",
          "post_date": "2021-07-19T13:52:27.640000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/shangweichen\" target=\"_blank\">@shangweichen</a> !!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1389534,
      "author_name": "Arnab Ray Chaudhuri",
      "author_url": "",
      "post_date": "2021-07-15T19:41:14.060000",
      "content": "<p>Thank you very much for saving so much time and space.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1390195,
          "author_name": "Jonathan Besomi",
          "author_url": "",
          "post_date": "2021-07-16T12:46:57.323000",
          "content": "<p>Glad the dataset can be useful! All the best for the competition! 👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1388851,
      "author_name": "laxman kusuma",
      "author_url": "",
      "post_date": "2021-07-15T09:05:29.050000",
      "content": "<p>Thanks for sharing …. magnificent reduce…</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1390204,
          "author_name": "Jonathan Besomi",
          "author_url": "",
          "post_date": "2021-07-16T12:55:20.143000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/laxmankusuma\" target=\"_blank\">@laxmankusuma</a> !! All the best for this competition and the future ones!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1388074,
      "author_name": "tharun_01",
      "author_url": "",
      "post_date": "2021-07-14T16:22:29.750000",
      "content": "<p>Firstly, many thanks for your help. <br>\nI just had a beginner question, there are 4 subfolders with lots of images in each, for one data point or one patient, which one use for model training or which images to choose for model training</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1388084,
          "author_name": "Jonathan Besomi",
          "author_url": "",
          "post_date": "2021-07-14T16:31:05.223000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/tharun2001\" target=\"_blank\">@tharun2001</a>!</p>\n<p><strong>Short answer</strong>: use all images you can for training.</p>\n<p><strong>Long answer</strong>: this is a binary classification task at the \"case\" level. For each \"case\", there are 4 folders and each folder contains multiple scans (of different types, you can <a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/data\" target=\"_blank\">read more in the data tab here</a>). Note also that the same folder structure is also available for the test part. Therefore, it's suggested to model the problem so that we can exploit all data available, for instance, making 3D reconstructions of the scans.</p>\n<p>Hope it helps. Good luck!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1547031,
      "author_name": "Tanmay Gupta",
      "author_url": "",
      "post_date": "2021-10-16T16:27:27.277000",
      "content": "<p>Thank you so much for doing this and sharing it with everyone! Really made my life easier!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1540781,
      "author_name": "Kwikik",
      "author_url": "",
      "post_date": "2021-10-10T22:37:27.200000",
      "content": "<p>\"data = (data * 255).astype(np.uint8)\" why using \"uint8\" dtype since decom image take value between 0 and 4096 ? dont we lose important information ? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1524936,
      "author_name": "Taiyo EG",
      "author_url": "",
      "post_date": "2021-09-27T02:26:15.343000",
      "content": "<p>Thanks for the great data.</p>\n<p>I'm trying to insert a directory structure data into kaggle's dataset instead of a file, but I can't do it.<br>\nHow did you manage to insert the data in a directory structure?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1507500,
      "author_name": "MikaYu",
      "author_url": "",
      "post_date": "2021-09-09T09:01:15.880000",
      "content": "<p>Thank you for sharing!  It's a very useful to use by Colaboratory.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1506049,
      "author_name": "alckasoc",
      "author_url": "",
      "post_date": "2021-09-07T19:14:08.987000",
      "content": "<p>Thanks for the dataset! Do you have the full script for converting everything to PNG?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1503215,
      "author_name": "Md Zarif Ul Alam",
      "author_url": "",
      "post_date": "2021-09-05T06:36:18.467000",
      "content": "<p>Thanks for sharing, helps with the training time :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1499113,
      "author_name": "Ayushman Buragohain",
      "author_url": "",
      "post_date": "2021-09-01T12:55:26.383000",
      "content": "<p>How did you filter out the empty DICOM images ? Could you share the script that you used to generate the dataset ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1485498,
      "author_name": "DeeplyAbstract",
      "author_url": "",
      "post_date": "2021-08-22T06:41:55.600000",
      "content": "<p>How does this code actually work out deep down? I'm curious.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1464832,
      "author_name": "Daniel Chen",
      "author_url": "",
      "post_date": "2021-08-10T18:18:42.247000",
      "content": "<p>How would I write the result <code>data</code> to a <code>.png</code> file?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1464996,
          "author_name": "Daniel Chen",
          "author_url": "",
          "post_date": "2021-08-10T19:46:12.593000",
          "content": "<pre><code>def process_dicom(path, outpath):\n    dicom = pydicom.read_file(path)\n    data = apply_voi_lut(dicom.pixel_array, dicom)\n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n\n    height = len(data)\n    width = len(data[0])\n\n    pixels_out = []\n    for row in data:\n        pixels_out.extend(row)\n\n    image_out = Image.new('L', (width, height))\n    image_out.save(outpath)\n</code></pre>\n<p>Tried something like this, I think the <code>mode</code> code is wrong! I am just getting a black image as output.  'L' is an 8-bit greyscale mode, see <a href=\"https://pillow.readthedocs.io/en/stable/handbook/concepts.html#concept-modes\" target=\"_blank\">https://pillow.readthedocs.io/en/stable/handbook/concepts.html#concept-modes</a>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1464998,
          "author_name": "Daniel Chen",
          "author_url": "",
          "post_date": "2021-08-10T19:46:38.560000",
          "content": "<p>was following this example: <a href=\"https://stackoverflow.com/questions/30943966/how-can-i-create-a-png-image-file-from-a-list-of-pixel-values-in-python\" target=\"_blank\">https://stackoverflow.com/questions/30943966/how-can-i-create-a-png-image-file-from-a-list-of-pixel-values-in-python</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1465015,
          "author_name": "Daniel Chen",
          "author_url": "",
          "post_date": "2021-08-10T20:01:26.530000",
          "content": "<p>Hah, forgot to do</p>\n<pre><code>    image_out.putdata(pixels_out)\n</code></pre>\n<p>After creating the PIL object! It works now!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1464692,
      "author_name": "Daniel Chen",
      "author_url": "",
      "post_date": "2021-08-10T17:03:25.943000",
      "content": "<p>Apparently RE: this link <a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/263612#1464675\" target=\"_blank\">https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/263612#1464675</a></p>\n<p>We can't actually use the PNG dataset when submitting we have to convert from DICOM to PNG ourselves.</p>\n<p>Regardless this is very helpful, thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1461562,
      "author_name": "Vadzim Tsitko",
      "author_url": "",
      "post_date": "2021-08-09T13:09:33.177000",
      "content": "<p>Cool! <br>\nBut it is necessary to understand, that DICOM contains not only an image, but also different additional data. <a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/252942\" target=\"_blank\">Here</a> Peter showed how to extract it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1460162,
      "author_name": "Fowaad",
      "author_url": "",
      "post_date": "2021-08-08T17:56:04.167000",
      "content": "<p>Awesome!!, Thank you so much for sharing this! :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1407313,
      "author_name": "Sougata",
      "author_url": "",
      "post_date": "2021-08-01T18:08:22.900000",
      "content": "<p>Cool. Thats really helpful</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1403601,
      "author_name": "Rohaan Nadeem",
      "author_url": "",
      "post_date": "2021-07-29T09:11:54.803000",
      "content": "<p>Hey,</p>\n<p>Thank you for the effort but I am curious, doesn't this conversion compromise the Meta-Data? Also, our inputs would be DICOM images, right?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1403028,
      "author_name": "Mohd Faraz Shaikh",
      "author_url": "",
      "post_date": "2021-07-28T18:12:58.263000",
      "content": "<p>Hii, thank you for sharing your work with us. Can you also tell me how do I import the reduced dataset in google colab notebooks?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1402521,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-28T08:57:15.930000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1401825,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-27T15:32:43.177000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1398191,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-23T20:15:10.140000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1395919,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-21T16:06:00.257000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1395298,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-21T05:46:41.627000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1395281,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-21T05:23:05.650000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1394957,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-20T18:49:56.077000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1394510,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-20T12:21:57.700000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1391993,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-18T09:21:30.020000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1391000,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-17T08:22:26.017000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1393258,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-19T13:52:11.523000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1390999,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-17T08:22:20.667000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1389165,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-15T13:44:39.600000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1390199,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-16T12:51:22.003000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1392583,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-18T20:47:35.177000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1402406,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-28T07:04:35.300000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1402463,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-28T08:15:14.090000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1407359,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-08-01T19:08:29.263000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1389162,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-15T13:42:21.547000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1390202,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-16T12:54:44.793000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1391999,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-18T09:28:22.420000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1388350,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-14T21:06:12.237000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1390210,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-16T13:00:52.087000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1485282,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-22T00:15:39.883000",
      "content": "",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1400510,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-26T11:02:11.120000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1395003,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-20T19:46:17.440000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1393109,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-19T11:40:30.547000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1393244,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-19T13:44:42.103000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1393823,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-20T00:09:33.667000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1392050,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-18T10:28:23.657000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1393252,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-19T13:51:00.733000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1389521,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-15T19:22:43.463000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1388414,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-14T23:28:34.737000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1388747,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-15T07:36:08.020000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1389579,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-15T20:51:14.420000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1390205,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-16T12:59:31.167000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1538355,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-10-08T10:34:06.897000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1535553,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-10-06T00:37:46.337000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1508350,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-09-10T06:59:43.050000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1507707,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-09-09T13:03:12.023000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1490981,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-26T04:15:24.700000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1485310,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-22T01:42:25.983000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1477503,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-17T13:59:11.910000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1473915,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-15T19:17:01.260000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1473423,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-15T15:05:44.153000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1473137,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-15T11:41:05.720000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1457486,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-07T12:40:18.997000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1443214,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-04T02:12:53.270000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1400843,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-26T16:25:57.900000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1400319,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-26T07:54:06.937000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1400037,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-26T00:54:52.077000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1397491,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-23T08:41:11.983000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1396288,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-22T01:43:03.203000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1395033,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-20T20:27:02.627000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1393896,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-20T02:25:59.653000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1391078,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-17T09:12:21.437000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1388021": "Preprocessing DICOM files is generally quite computational-intensive and slow, especially considering the large dataset size (~128 GB). For this reason, I've fired up a GCP VM and transformed the data from DICOM to PNG for you. The dataset file size has been reduced from 128 GB to 5.2 GB. \n\n**Notes**\n\n - Images sizes have been kept as the original ones. \n - To further reduce file size, all empty DICOM images files have not been included in the dataset, you can easily spot them by looking at the images sequences (`Image-X.png`). Note that, because of that, some entire folders (such as `train/00109/FLAIR`) are not included in this dataset (as they contain empty images only).\n - The structure of the dataset is not changed. I.e all _.dcm_ files are now _.png_ files.\n\n**Dataset link**\n\n - [RSNA-MICCAI-PNG-Dataset](https://www.kaggle.com/jonathanbesomi/rsna-miccai-png) [5.2 GB]\n\n**Script**\n\nFor each image, run: (Adapted from @tanlikesmath [work](https://www.kaggle.com/tanlikesmath/brain-tumor-radiogenomic-classification-eda))\n\n```\ndicom = pydicom.read_file(path)\ndata = apply_voi_lut(dicom.pixel_array, dicom)\nif dicom.PhotometricInterpretation == \"MONOCHROME1\":\n    data = np.amax(data) - data\ndata = data - np.min(data)\ndata = data / np.max(data)\ndata = (data * 255).astype(np.uint8)\n```\n\n**More**\n\nIf you need other formats, you want a version with resized images or anything else, just ask 👍\n\nThank you for reading, happy kaggling to all!",
    "1395928": "Greate work.\n\nCould you please elaborate on the decision to normalize an images the following way?\n```\ndata = data - np.min(data)\ndata = data / np.max(data)\ndata = (data * 255).astype(np.uint8)\n```",
    "1389541": "Thanks for sharing",
    "1392053": "Thank you for sharing this..",
    "1471361": "Why do you use this normalization formula?\n`data = data - np.min(data)`\n  `data = data / np.max(data)`\nI checked the original Dicom file, some local minimal values are not equal to the global minimum value, so is the maximum value. The global minimal value is -32768.0 and the global maximum value is 65535.0 (I only checked the converted dicom files). But the minimum value of \"Image-118.dcm\" is -32754.13895939086, and the maximum value of \"Image-118.dcm\" is 16424.83312182741. Your formula threats the -32754 as 0 and 16424 as 255, I think it's wrong.\nI am considering to use this formula:\n`global_min, global_max = -32768.0, 65535.0`\n`data = data - global_min`\n    `data = data / global_max`\nThe converted images look similar, but they have different pixel values.",
    "1389296": "Thanks a lot!\nBut why using min-max normalize instead of normalize them using their full range (i.e. 0 to 2 ** bit_stored)??",
    "1388314": "Thanks for sharing! ~~Just for those who don't know the size of the images. I just download them and check. The size of the png images is 512 x 512.~~",
    "1388037": "Wow, this is amazing @jonathanbesomi! My computer has an SSD, so it's limited when it comes to space. Your solution will make this dataset way more accessible to people like me! Thanks a lot!",
    "1502064": "Thanks a lot for sharing, this is very kind of you.",
    "1392161": "Great work Thanks for sharing !",
    "1391397": "Thanks for sharing ! Extremely useful",
    "1390308": "Pretty amazing work, thank you for sharing with community.",
    "1389534": "Thank you very much for saving so much time and space.",
    "1388851": "Thanks for sharing .... magnificent reduce...",
    "1388074": "Firstly, many thanks for your help. \nI just had a beginner question, there are 4 subfolders with lots of images in each, for one data point or one patient, which one use for model training or which images to choose for model training",
    "1547031": "Thank you so much for doing this and sharing it with everyone! Really made my life easier!",
    "1540781": "\"data = (data * 255).astype(np.uint8)\" why using \"uint8\" dtype since decom image take value between 0 and 4096 ? dont we lose important information ? ",
    "1524936": "Thanks for the great data.\n\nI'm trying to insert a directory structure data into kaggle's dataset instead of a file, but I can't do it.\nHow did you manage to insert the data in a directory structure?\n\n\n",
    "1507500": "Thank you for sharing!  It's a very useful to use by Colaboratory.",
    "1506049": "Thanks for the dataset! Do you have the full script for converting everything to PNG?",
    "1503215": "Thanks for sharing, helps with the training time :)",
    "1499113": "How did you filter out the empty DICOM images ? Could you share the script that you used to generate the dataset ?",
    "1485498": "How does this code actually work out deep down? I'm curious.",
    "1464832": "How would I write the result `data` to a `.png` file?",
    "1464692": "Apparently RE: this link https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/263612#1464675\n\nWe can't actually use the PNG dataset when submitting we have to convert from DICOM to PNG ourselves.\n\nRegardless this is very helpful, thank you!",
    "1461562": "Cool! \nBut it is necessary to understand, that DICOM contains not only an image, but also different additional data. [Here](https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/252942) Peter showed how to extract it.",
    "1460162": "Awesome!!, Thank you so much for sharing this! :)",
    "1407313": "Cool. Thats really helpful",
    "1403601": "Hey,\n\nThank you for the effort but I am curious, doesn't this conversion compromise the Meta-Data? Also, our inputs would be DICOM images, right?",
    "1403028": "Hii, thank you for sharing your work with us. Can you also tell me how do I import the reduced dataset in google colab notebooks?",
    "1402521": "Excellent!",
    "1401825": "Thank you for sharing great & helpful information!",
    "1398191": "Thank you to share your knowledge! Great work!! ",
    "1395919": "Is there any link to the data where just format changed from dicom to png/jpeg, keeping other things exactly same, like not removing those empty image files etc ?\nIt would be a great help, thanks.",
    "1395298": "Thank for your help!!!\nCan you share more sizes? More sizes will be very interesting!",
    "1395281": "Thanks for sharing I was looking for this and this will be really helpful for me.",
    "1394957": "Great work.",
    "1394510": "Sincerely Thanks for sharing this it will really help us in training our model",
    "1391993": "Glad to see something I was looking for quite a while",
    "1391000": "That is brilliant. Thank you for time and energy-saving tip, John!",
    "1390999": "That is brilliant. Thank you for time and energy-saving tip, John!",
    "1389165": "Thanks for the transformation! any possibility of having the data in .nii.gz format? I am getting 3 different volumes for each image when applying dcm2nii",
    "1389162": "You saved my internet. And my bill. Thanks!",
    "1388350": "You have a place reserved in Heaven.",
    "1485282": "Request to add train label file too😄. Thanks for your efforts you made to convert the file in .png file. Because of your efforts, it's now practical possible for me to work with. ",
    "1400510": "",
    "1395003": "",
    "1393109": "",
    "1392050": "Thanks for the sharing. As the target is radiomics, is precision lost (from 16 bits to 8bits) can damage subtitles textural patterns ? :s ",
    "1389521": "",
    "1388414": "",
    "1538355": "Thanks a lot!",
    "1535553": "Thanks for sharing!",
    "1508350": "Thank you very much!",
    "1507707": "Thank you for sharing :) ",
    "1490981": "Thank you for sharing!",
    "1485310": "Thank you  for sharing!!",
    "1477503": "Awesome work! Thank you for sharing!",
    "1473915": "thanks for sharing great work",
    "1473423": "Thank you for sharing!",
    "1473137": "Thanks for sharing! very helpful!",
    "1457486": "thank you so much\n",
    "1443214": "Thank you you saved my disk.... :)",
    "1400843": "Thanks a lot !",
    "1400319": "Thank you for sharing :)",
    "1400037": "Thank you Jonathan :D",
    "1397491": "Great and Thanks for sharing this .\n",
    "1396288": "Great! Thanks!",
    "1395033": "Thank you so much for sharing.",
    "1393896": "Thank you for such a nice work :)",
    "1391078": "Thank you for sharing."
  }
}