{
  "id": 268393,
  "title": "Normalisation",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/268393",
  "author_name": "",
  "post_date": "2021-08-27T07:53:18.817749200Z",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I've read a few discussions on how to normalise the dataset. I agree with the consensus that for a given MRI type, normalisation should be implemented over all slices as opposed to normalising each slice separately, since per slice normalisation would lead to significant discontinuities.</p>\n<p>However, this got me thinking that it may be desirable to normalise across the entire training data set (for each MRI type). For instance, suppose that one tumour type is generally more dense than the other (this may or may not be the case); then we would lose this information when normalising each patient's data separately. This could affect a model's ability to distinguish between tumour types. Has anyone found this to make a difference? </p>",
  "messages": [
    {
      "id": "1492489",
      "postDate": "08/27/2021 07:53:18",
      "content": "<p>I've read a few discussions on how to normalise the dataset. I agree with the consensus that for a given MRI type, normalisation should be implemented over all slices as opposed to normalising each slice separately, since per slice normalisation would lead to significant discontinuities.</p>\n<p>However, this got me thinking that it may be desirable to normalise across the entire training data set (for each MRI type). For instance, suppose that one tumour type is generally more dense than the other (this may or may not be the case); then we would lose this information when normalising each patient's data separately. This could affect a model's ability to distinguish between tumour types. Has anyone found this to make a difference? </p>",
      "rawMarkdown": "I've read a few discussions on how to normalise the dataset. I agree with the consensus that for a given MRI type, normalisation should be implemented over all slices as opposed to normalising each slice separately, since per slice normalisation would lead to significant discontinuities.\n\nHowever, this got me thinking that it may be desirable to normalise across the entire training data set (for each MRI type). For instance, suppose that one tumour type is generally more dense than the other (this may or may not be the case); then we would lose this information when normalising each patient's data separately. This could affect a model's ability to distinguish between tumour types. Has anyone found this to make a difference?",
      "votes": null
    },
    {
      "id": "1492814",
      "postDate": "08/27/2021 12:49:49",
      "content": "<p>Not sure if normalization is needed. I think each slice has different brightness. This is 00000 FLAIR:<br>\n![<a href=\"https://media.discordapp.net/attachments/696664023365713933/880795727561916446/results___11_0.png\" target=\"_blank\">https://media.discordapp.net/attachments/696664023365713933/880795727561916446/results___11_0.png</a>]</p>\n<p>On the y-axis you can see the vertical layers of brightness.</p>",
      "rawMarkdown": "Not sure if normalization is needed. I think each slice has different brightness. This is 00000 FLAIR:\n![https://media.discordapp.net/attachments/696664023365713933/880795727561916446/results___11_0.png]\n\nOn the y-axis you can see the vertical layers of brightness.",
      "votes": null
    },
    {
      "id": "1493429",
      "postDate": "08/27/2021 21:54:54",
      "content": "<p>Since MR images are mostly more than 8 bit .. some amount of normalization automagically happens when they're exported to JPG. This is usually done with simple linear normalization. Like pydicom or matplotlib does, or something like this ..</p>\n<pre><code>pixels = pixels - np.min(pixels)\npixels = pixels / np.max(pixels)\npixels = (pixels * 255).astype(np.uint8)\n</code></pre>\n<p>The challenge with this approach is that it changes the pixel distribution and undoes the inherent contrast in the raw pixels in a lot of studies. Especially on mages which have had the bone removed, yet small remnants of bone still exist (like these). It works fine for some, and bad for others.</p>\n<p>Maintaining the original pixel distribution that exists at the 16 bit level should be priority number one. MR's are designed to demonstrate subtle contrast between tissue types. Anything that changes contrast (normalization, histogram equalization, CLAHE etc), should be used sparingly or not at all. </p>\n<p>This won't be easy because of varying machines, protocols and post-processing. Also, I think some of the studies were reprocessed after having bone removed and some were not. The DICOM default window width and window center settings are unreliable in a lot of studies.</p>\n<p>Ideally, manual window width and center settings (preset windows) should be used to standardize across studies.</p>",
      "rawMarkdown": "Since MR images are mostly more than 8 bit .. some amount of normalization automagically happens when they're exported to JPG. This is usually done with simple linear normalization. Like pydicom or matplotlib does, or something like this ..\n```\npixels = pixels - np.min(pixels)\npixels = pixels / np.max(pixels)\npixels = (pixels * 255).astype(np.uint8)\n```\nThe challenge with this approach is that it changes the pixel distribution and undoes the inherent contrast in the raw pixels in a lot of studies. Especially on mages which have had the bone removed, yet small remnants of bone still exist (like these). It works fine for some, and bad for others.\n\nMaintaining the original pixel distribution that exists at the 16 bit level should be priority number one. MR's are designed to demonstrate subtle contrast between tissue types. Anything that changes contrast (normalization, histogram equalization, CLAHE etc), should be used sparingly or not at all. \n\nThis won't be easy because of varying machines, protocols and post-processing. Also, I think some of the studies were reprocessed after having bone removed and some were not. The DICOM default window width and window center settings are unreliable in a lot of studies.\n\nIdeally, manual window width and center settings (preset windows) should be used to standardize across studies.",
      "votes": null
    },
    {
      "id": "1493838",
      "postDate": "08/28/2021 07:36:10",
      "content": "<p>Thanks a lot for the detailed reply. </p>\n<p>Also, I've now noticed that MR images are not standardised like CT scans. So normalisation across all patient data is probably not a great idea.</p>",
      "rawMarkdown": "Thanks a lot for the detailed reply. \n\nAlso, I've now noticed that MR images are not standardised like CT scans. So normalisation across all patient data is probably not a great idea.",
      "votes": null
    },
    {
      "id": "1495228",
      "postDate": "08/29/2021 11:53:26",
      "content": "<p>I found that many images has np.min(pixels) == np.max(pixels) == np.mean(pixels). In such cases the normalization doesn't work. How to deal with these images?</p>\n<p>For example,<br>\na = load_dicom_images_3d('00043')<br>\nprint(a.shape,a.min(),a.max(),a.mean())</p>\n<p>(512, 512) 32767.0 32767.0 32767.0<br>\n…</p>",
      "rawMarkdown": "I found that many images has np.min(pixels) == np.max(pixels) == np.mean(pixels). In such cases the normalization doesn't work. How to deal with these images?\n\nFor example,\na = load_dicom_images_3d('00043')\nprint(a.shape,a.min(),a.max(),a.mean())\n\n(512, 512) 32767.0 32767.0 32767.0\n...",
      "votes": null
    },
    {
      "id": "1495242",
      "postDate": "08/29/2021 12:05:30",
      "content": "<p>These images just represent empty space. You can use something like:</p>\n<pre><code>if a.max() - a.min() == 0:  \n    a = 0\nelse:\n    a = normalise(a)\n</code></pre>\n<p>The question as to what to do with this zero data is up to you though. You can keep it in the dataset, or not. Not sure what will work best.</p>",
      "rawMarkdown": "These images just represent empty space. You can use something like:\n```\nif a.max() - a.min() == 0:  \n    a = 0\nelse:\n    a = normalise(a)\n```\nThe question as to what to do with this zero data is up to you though. You can keep it in the dataset, or not. Not sure what will work best.",
      "votes": null
    },
    {
      "id": "1495262",
      "postDate": "08/29/2021 12:21:08",
      "content": "<p>JJ, thanks for your reply!<br>\nI guessed that it's just an empty images 5 minutes after posting..) You're right, it's not so obvious just to throw them out.</p>",
      "rawMarkdown": "JJ, thanks for your reply!\nI guessed that it's just an empty images 5 minutes after posting..) You're right, it's not so obvious just to throw them out.",
      "votes": null
    },
    {
      "id": "1496985",
      "postDate": "08/30/2021 19:39:27",
      "content": "<p>Did you use apply_voi_lut? If yes, try without.</p>",
      "rawMarkdown": "Did you use apply_voi_lut? If yes, try without.",
      "votes": null
    },
    {
      "id": "1500322",
      "postDate": "09/02/2021 09:32:20",
      "content": "<p>The question is important. I have also wondered it a littlebit and has done implemented some thinking of it to a notebook <a href=\"https://www.kaggle.com/experienceinai/aivo-vii-nopea\" target=\"_blank\">https://www.kaggle.com/experienceinai/aivo-vii-nopea</a>. It includes function analysoi_yhden_kansion_kuvapinon_piirteet, which looks in to the \"stack of consequtive images\" and its characteristics in time domain. This time domain curve varies a lot between individual measurements (in notebook the green curve in picture illustrating \"interesting images\"). This curve is some form of the measurement magnetic pulse. My intuition is its shape should be taken into account when making normalization. Has anyone done it yet? Can the \"baseline\" form of the measurement magnetic pulse be found somewhere? E.g. from the DICOM metadata?</p>",
      "rawMarkdown": "The question is important. I have also wondered it a littlebit and has done implemented some thinking of it to a notebook https://www.kaggle.com/experienceinai/aivo-vii-nopea. It includes function analysoi_yhden_kansion_kuvapinon_piirteet, which looks in to the \"stack of consequtive images\" and its characteristics in time domain. This time domain curve varies a lot between individual measurements (in notebook the green curve in picture illustrating \"interesting images\"). This curve is some form of the measurement magnetic pulse. My intuition is its shape should be taken into account when making normalization. Has anyone done it yet? Can the \"baseline\" form of the measurement magnetic pulse be found somewhere? E.g. from the DICOM metadata?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1492814,
      "author_name": "aleksandrkruchinin",
      "author_url": "",
      "post_date": "08/27/2021 12:49:49",
      "content": "<p>Not sure if normalization is needed. I think each slice has different brightness. This is 00000 FLAIR:<br>\n![<a href=\"https://media.discordapp.net/attachments/696664023365713933/880795727561916446/results___11_0.png\" target=\"_blank\">https://media.discordapp.net/attachments/696664023365713933/880795727561916446/results___11_0.png</a>]</p>\n<p>On the y-axis you can see the vertical layers of brightness.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1496985,
          "author_name": "bpetrb",
          "author_url": "",
          "post_date": "08/30/2021 19:39:27",
          "content": "<p>Did you use apply_voi_lut? If yes, try without.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1493429,
      "author_name": "davidbroberts",
      "author_url": "",
      "post_date": "08/27/2021 21:54:54",
      "content": "<p>Since MR images are mostly more than 8 bit .. some amount of normalization automagically happens when they're exported to JPG. This is usually done with simple linear normalization. Like pydicom or matplotlib does, or something like this ..</p>\n<pre><code>pixels = pixels - np.min(pixels)\npixels = pixels / np.max(pixels)\npixels = (pixels * 255).astype(np.uint8)\n</code></pre>\n<p>The challenge with this approach is that it changes the pixel distribution and undoes the inherent contrast in the raw pixels in a lot of studies. Especially on mages which have had the bone removed, yet small remnants of bone still exist (like these). It works fine for some, and bad for others.</p>\n<p>Maintaining the original pixel distribution that exists at the 16 bit level should be priority number one. MR's are designed to demonstrate subtle contrast between tissue types. Anything that changes contrast (normalization, histogram equalization, CLAHE etc), should be used sparingly or not at all. </p>\n<p>This won't be easy because of varying machines, protocols and post-processing. Also, I think some of the studies were reprocessed after having bone removed and some were not. The DICOM default window width and window center settings are unreliable in a lot of studies.</p>\n<p>Ideally, manual window width and center settings (preset windows) should be used to standardize across studies.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1493838,
          "author_name": "jjbennett",
          "author_url": "",
          "post_date": "08/28/2021 07:36:10",
          "content": "<p>Thanks a lot for the detailed reply. </p>\n<p>Also, I've now noticed that MR images are not standardised like CT scans. So normalisation across all patient data is probably not a great idea.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1495228,
          "author_name": "polishch",
          "author_url": "",
          "post_date": "08/29/2021 11:53:26",
          "content": "<p>I found that many images has np.min(pixels) == np.max(pixels) == np.mean(pixels). In such cases the normalization doesn't work. How to deal with these images?</p>\n<p>For example,<br>\na = load_dicom_images_3d('00043')<br>\nprint(a.shape,a.min(),a.max(),a.mean())</p>\n<p>(512, 512) 32767.0 32767.0 32767.0<br>\n…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1495242,
          "author_name": "jjbennett",
          "author_url": "",
          "post_date": "08/29/2021 12:05:30",
          "content": "<p>These images just represent empty space. You can use something like:</p>\n<pre><code>if a.max() - a.min() == 0:  \n    a = 0\nelse:\n    a = normalise(a)\n</code></pre>\n<p>The question as to what to do with this zero data is up to you though. You can keep it in the dataset, or not. Not sure what will work best.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1495262,
              "author_name": "polishch",
              "author_url": "",
              "post_date": "08/29/2021 12:21:08",
              "content": "<p>JJ, thanks for your reply!<br>\nI guessed that it's just an empty images 5 minutes after posting..) You're right, it's not so obvious just to throw them out.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1500322,
      "author_name": "experienceinai",
      "author_url": "",
      "post_date": "09/02/2021 09:32:20",
      "content": "<p>The question is important. I have also wondered it a littlebit and has done implemented some thinking of it to a notebook <a href=\"https://www.kaggle.com/experienceinai/aivo-vii-nopea\" target=\"_blank\">https://www.kaggle.com/experienceinai/aivo-vii-nopea</a>. It includes function analysoi_yhden_kansion_kuvapinon_piirteet, which looks in to the \"stack of consequtive images\" and its characteristics in time domain. This time domain curve varies a lot between individual measurements (in notebook the green curve in picture illustrating \"interesting images\"). This curve is some form of the measurement magnetic pulse. My intuition is its shape should be taken into account when making normalization. Has anyone done it yet? Can the \"baseline\" form of the measurement magnetic pulse be found somewhere? E.g. from the DICOM metadata?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1492489": "I've read a few discussions on how to normalise the dataset. I agree with the consensus that for a given MRI type, normalisation should be implemented over all slices as opposed to normalising each slice separately, since per slice normalisation would lead to significant discontinuities.\n\nHowever, this got me thinking that it may be desirable to normalise across the entire training data set (for each MRI type). For instance, suppose that one tumour type is generally more dense than the other (this may or may not be the case); then we would lose this information when normalising each patient's data separately. This could affect a model's ability to distinguish between tumour types. Has anyone found this to make a difference?",
    "1492814": "Not sure if normalization is needed. I think each slice has different brightness. This is 00000 FLAIR:\n![https://media.discordapp.net/attachments/696664023365713933/880795727561916446/results___11_0.png]\n\nOn the y-axis you can see the vertical layers of brightness.",
    "1493429": "Since MR images are mostly more than 8 bit .. some amount of normalization automagically happens when they're exported to JPG. This is usually done with simple linear normalization. Like pydicom or matplotlib does, or something like this ..\n```\npixels = pixels - np.min(pixels)\npixels = pixels / np.max(pixels)\npixels = (pixels * 255).astype(np.uint8)\n```\nThe challenge with this approach is that it changes the pixel distribution and undoes the inherent contrast in the raw pixels in a lot of studies. Especially on mages which have had the bone removed, yet small remnants of bone still exist (like these). It works fine for some, and bad for others.\n\nMaintaining the original pixel distribution that exists at the 16 bit level should be priority number one. MR's are designed to demonstrate subtle contrast between tissue types. Anything that changes contrast (normalization, histogram equalization, CLAHE etc), should be used sparingly or not at all. \n\nThis won't be easy because of varying machines, protocols and post-processing. Also, I think some of the studies were reprocessed after having bone removed and some were not. The DICOM default window width and window center settings are unreliable in a lot of studies.\n\nIdeally, manual window width and center settings (preset windows) should be used to standardize across studies.",
    "1493838": "Thanks a lot for the detailed reply. \n\nAlso, I've now noticed that MR images are not standardised like CT scans. So normalisation across all patient data is probably not a great idea.",
    "1495228": "I found that many images has np.min(pixels) == np.max(pixels) == np.mean(pixels). In such cases the normalization doesn't work. How to deal with these images?\n\nFor example,\na = load_dicom_images_3d('00043')\nprint(a.shape,a.min(),a.max(),a.mean())\n\n(512, 512) 32767.0 32767.0 32767.0\n...",
    "1495242": "These images just represent empty space. You can use something like:\n```\nif a.max() - a.min() == 0:  \n    a = 0\nelse:\n    a = normalise(a)\n```\nThe question as to what to do with this zero data is up to you though. You can keep it in the dataset, or not. Not sure what will work best.",
    "1495262": "JJ, thanks for your reply!\nI guessed that it's just an empty images 5 minutes after posting..) You're right, it's not so obvious just to throw them out.",
    "1496985": "Did you use apply_voi_lut? If yes, try without.",
    "1500322": "The question is important. I have also wondered it a littlebit and has done implemented some thinking of it to a notebook https://www.kaggle.com/experienceinai/aivo-vii-nopea. It includes function analysoi_yhden_kansion_kuvapinon_piirteet, which looks in to the \"stack of consequtive images\" and its characteristics in time domain. This time domain curve varies a lot between individual measurements (in notebook the green curve in picture illustrating \"interesting images\"). This curve is some form of the measurement magnetic pulse. My intuition is its shape should be taken into account when making normalization. Has anyone done it yet? Can the \"baseline\" form of the measurement magnetic pulse be found somewhere? E.g. from the DICOM metadata?"
  },
  "source": "meta"
}