{
  "id": 220182,
  "title": "Image Normalization:   Is it worthwhile?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/220182",
  "author_name": "",
  "post_date": "2021-02-17T16:10:50.158913Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I did a quick experiment to calculate the mean and standard deviation of the test and training images for each channel.   The results are:</p>\n<pre><code>test\nfor afa5e8098 stats are (array([158, 120, 175]), array([30, 39, 22]))\nfor b9a3865fc stats are (array([162, 114, 182]), array([38, 44, 26]))\nfor c68fe75ea stats are (array([165, 124, 177]), array([27, 36, 20]))\nfor b2dc8411c stats are (array([164, 126, 183]), array([43, 57, 30]))\nfor 26dc41664 stats are (array([143,  81, 155]), array([29, 39, 22]))\ntrain\nfor 095bf7a1f stats are (array([146,  84, 151]), array([24, 33, 21]))\nfor e79de561c stats are (array([147, 101, 165]), array([32, 43, 23]))\nfor 0486052bb stats are (array([163, 128, 185]), array([40, 50, 28]))\nfor aaa6a05cc stats are (array([148,  96, 170]), array([44, 50, 32]))\nfor 2f6ecfcdf stats are (array([204, 188, 211]), array([24, 35, 16]))\nfor 54f2eec69 stats are (array([146,  90, 161]), array([25, 35, 18]))\nfor 1e2425f28 stats are (array([156,  82, 155]), array([25, 40, 23]))\nfor cb2d976f4 stats are (array([190, 153, 200]), array([29, 44, 21]))\n</code></pre>\n<p>My question is:   Is there enough variation here to affect the predictions?   If the answer is \"yes\", then what I'd do is:</p>\n<ol>\n<li>Calculate the per-channel means and standard deviations for all the \"test\" and \"train\" images, weighting the contribution of each image by the number of pixels in the image.</li>\n<li>For Training, normalize all the \"train\" images per-channel by subtracting the mean and dividing by the std deviation.</li>\n<li>For Submission, normalize all the submission images per-channel by subtracting the mean and dividing by the std deviation.</li>\n</ol>\n<p>Again, the question is:   Is there enough variation that this is worthwhile?</p>",
  "messages": [
    {
      "id": "1206937",
      "postDate": "02/17/2021 16:10:50",
      "content": "<p>I did a quick experiment to calculate the mean and standard deviation of the test and training images for each channel.   The results are:</p>\n<pre><code>test\nfor afa5e8098 stats are (array([158, 120, 175]), array([30, 39, 22]))\nfor b9a3865fc stats are (array([162, 114, 182]), array([38, 44, 26]))\nfor c68fe75ea stats are (array([165, 124, 177]), array([27, 36, 20]))\nfor b2dc8411c stats are (array([164, 126, 183]), array([43, 57, 30]))\nfor 26dc41664 stats are (array([143,  81, 155]), array([29, 39, 22]))\ntrain\nfor 095bf7a1f stats are (array([146,  84, 151]), array([24, 33, 21]))\nfor e79de561c stats are (array([147, 101, 165]), array([32, 43, 23]))\nfor 0486052bb stats are (array([163, 128, 185]), array([40, 50, 28]))\nfor aaa6a05cc stats are (array([148,  96, 170]), array([44, 50, 32]))\nfor 2f6ecfcdf stats are (array([204, 188, 211]), array([24, 35, 16]))\nfor 54f2eec69 stats are (array([146,  90, 161]), array([25, 35, 18]))\nfor 1e2425f28 stats are (array([156,  82, 155]), array([25, 40, 23]))\nfor cb2d976f4 stats are (array([190, 153, 200]), array([29, 44, 21]))\n</code></pre>\n<p>My question is:   Is there enough variation here to affect the predictions?   If the answer is \"yes\", then what I'd do is:</p>\n<ol>\n<li>Calculate the per-channel means and standard deviations for all the \"test\" and \"train\" images, weighting the contribution of each image by the number of pixels in the image.</li>\n<li>For Training, normalize all the \"train\" images per-channel by subtracting the mean and dividing by the std deviation.</li>\n<li>For Submission, normalize all the submission images per-channel by subtracting the mean and dividing by the std deviation.</li>\n</ol>\n<p>Again, the question is:   Is there enough variation that this is worthwhile?</p>",
      "rawMarkdown": "I did a quick experiment to calculate the mean and standard deviation of the test and training images for each channel.   The results are:\n```\ntest\nfor afa5e8098 stats are (array([158, 120, 175]), array([30, 39, 22]))\nfor b9a3865fc stats are (array([162, 114, 182]), array([38, 44, 26]))\nfor c68fe75ea stats are (array([165, 124, 177]), array([27, 36, 20]))\nfor b2dc8411c stats are (array([164, 126, 183]), array([43, 57, 30]))\nfor 26dc41664 stats are (array([143,  81, 155]), array([29, 39, 22]))\ntrain\nfor 095bf7a1f stats are (array([146,  84, 151]), array([24, 33, 21]))\nfor e79de561c stats are (array([147, 101, 165]), array([32, 43, 23]))\nfor 0486052bb stats are (array([163, 128, 185]), array([40, 50, 28]))\nfor aaa6a05cc stats are (array([148,  96, 170]), array([44, 50, 32]))\nfor 2f6ecfcdf stats are (array([204, 188, 211]), array([24, 35, 16]))\nfor 54f2eec69 stats are (array([146,  90, 161]), array([25, 35, 18]))\nfor 1e2425f28 stats are (array([156,  82, 155]), array([25, 40, 23]))\nfor cb2d976f4 stats are (array([190, 153, 200]), array([29, 44, 21]))\n```\nMy question is:   Is there enough variation here to affect the predictions?   If the answer is \"yes\", then what I'd do is:\n1.  Calculate the per-channel means and standard deviations for all the \"test\" and \"train\" images, weighting the contribution of each image by the number of pixels in the image.\n2.  For Training, normalize all the \"train\" images per-channel by subtracting the mean and dividing by the std deviation.\n3.  For Submission, normalize all the submission images per-channel by subtracting the mean and dividing by the std deviation.\n\nAgain, the question is:   Is there enough variation that this is worthwhile?",
      "votes": null
    },
    {
      "id": "1215615",
      "postDate": "02/23/2021 19:40:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/markalavin\" target=\"_blank\">@markalavin</a> Usually normalization leads to faster convergence. So it is always worthwhile to try it out.<br>\nIn my personal experience it does not always deliver on its promises but as said … try it out.</p>\n<p>Use the same normalization for Train, Test and Validation. Good luck.</p>",
      "rawMarkdown": "Hi @markalavin Usually normalization leads to faster convergence. So it is always worthwhile to try it out.\nIn my personal experience it does not always deliver on its promises but as said ... try it out.\n\nUse the same normalization for Train, Test and Validation. Good luck.",
      "votes": null
    },
    {
      "id": "1218015",
      "postDate": "02/25/2021 13:40:20",
      "content": "<p>My guess is that it also depends on the architecture. Network like EfficientNet includes batch normalization layers (though not as the first layer), so benefits may not be that large. I also wonder if doing per-channel pre-normalization would introduce additional errors, like a mismatch between values in different channels.</p>",
      "rawMarkdown": "My guess is that it also depends on the architecture. Network like EfficientNet includes batch normalization layers (though not as the first layer), so benefits may not be that large. I also wonder if doing per-channel pre-normalization would introduce additional errors, like a mismatch between values in different channels.",
      "votes": null
    },
    {
      "id": "1218116",
      "postDate": "02/25/2021 15:18:13",
      "content": "<p>I agree with you on the point 'mismatch'. It sounds reasonable.</p>",
      "rawMarkdown": "I agree with you on the point 'mismatch'. It sounds reasonable.",
      "votes": null
    },
    {
      "id": "1218167",
      "postDate": "02/25/2021 16:00:50",
      "content": "<p>Here's how I did the normalization, and what the results were:</p>\n<p>Deviating from the original description, I normalized each image independent of the others, i.e., I did not collect statistics for all the images together.   Rather, for each image, for each channel, I calculated the mean and standard deviation, then subtracted the per-channel means from channel's values and divided it by the per-channel standard deviation and multiplied the result by 255/6 (because I want +/- 3 stddev's to map into [0-255]).</p>\n<p>When I tried building a model with these normalized images, I found that the out-of-fold validation DICE score was very low compared to what I had seen with the unnormalized images, e.g., 0.6 instead of 0.9.   So, I just gave up on the normalization idea.   But I'd gladly try anything you could suggest.</p>",
      "rawMarkdown": "Here's how I did the normalization, and what the results were:\n\nDeviating from the original description, I normalized each image independent of the others, i.e., I did not collect statistics for all the images together.   Rather, for each image, for each channel, I calculated the mean and standard deviation, then subtracted the per-channel means from channel's values and divided it by the per-channel standard deviation and multiplied the result by 255/6 (because I want +/- 3 stddev's to map into [0-255]).\n\nWhen I tried building a model with these normalized images, I found that the out-of-fold validation DICE score was very low compared to what I had seen with the unnormalized images, e.g., 0.6 instead of 0.9.   So, I just gave up on the normalization idea.   But I'd gladly try anything you could suggest.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1215615,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "02/23/2021 19:40:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/markalavin\" target=\"_blank\">@markalavin</a> Usually normalization leads to faster convergence. So it is always worthwhile to try it out.<br>\nIn my personal experience it does not always deliver on its promises but as said … try it out.</p>\n<p>Use the same normalization for Train, Test and Validation. Good luck.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1218015,
      "author_name": "gdonchyts",
      "author_url": "",
      "post_date": "02/25/2021 13:40:20",
      "content": "<p>My guess is that it also depends on the architecture. Network like EfficientNet includes batch normalization layers (though not as the first layer), so benefits may not be that large. I also wonder if doing per-channel pre-normalization would introduce additional errors, like a mismatch between values in different channels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1218116,
          "author_name": "louieshao",
          "author_url": "",
          "post_date": "02/25/2021 15:18:13",
          "content": "<p>I agree with you on the point 'mismatch'. It sounds reasonable.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1218167,
      "author_name": "markalavin",
      "author_url": "",
      "post_date": "02/25/2021 16:00:50",
      "content": "<p>Here's how I did the normalization, and what the results were:</p>\n<p>Deviating from the original description, I normalized each image independent of the others, i.e., I did not collect statistics for all the images together.   Rather, for each image, for each channel, I calculated the mean and standard deviation, then subtracted the per-channel means from channel's values and divided it by the per-channel standard deviation and multiplied the result by 255/6 (because I want +/- 3 stddev's to map into [0-255]).</p>\n<p>When I tried building a model with these normalized images, I found that the out-of-fold validation DICE score was very low compared to what I had seen with the unnormalized images, e.g., 0.6 instead of 0.9.   So, I just gave up on the normalization idea.   But I'd gladly try anything you could suggest.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1206937": "I did a quick experiment to calculate the mean and standard deviation of the test and training images for each channel.   The results are:\n```\ntest\nfor afa5e8098 stats are (array([158, 120, 175]), array([30, 39, 22]))\nfor b9a3865fc stats are (array([162, 114, 182]), array([38, 44, 26]))\nfor c68fe75ea stats are (array([165, 124, 177]), array([27, 36, 20]))\nfor b2dc8411c stats are (array([164, 126, 183]), array([43, 57, 30]))\nfor 26dc41664 stats are (array([143,  81, 155]), array([29, 39, 22]))\ntrain\nfor 095bf7a1f stats are (array([146,  84, 151]), array([24, 33, 21]))\nfor e79de561c stats are (array([147, 101, 165]), array([32, 43, 23]))\nfor 0486052bb stats are (array([163, 128, 185]), array([40, 50, 28]))\nfor aaa6a05cc stats are (array([148,  96, 170]), array([44, 50, 32]))\nfor 2f6ecfcdf stats are (array([204, 188, 211]), array([24, 35, 16]))\nfor 54f2eec69 stats are (array([146,  90, 161]), array([25, 35, 18]))\nfor 1e2425f28 stats are (array([156,  82, 155]), array([25, 40, 23]))\nfor cb2d976f4 stats are (array([190, 153, 200]), array([29, 44, 21]))\n```\nMy question is:   Is there enough variation here to affect the predictions?   If the answer is \"yes\", then what I'd do is:\n1.  Calculate the per-channel means and standard deviations for all the \"test\" and \"train\" images, weighting the contribution of each image by the number of pixels in the image.\n2.  For Training, normalize all the \"train\" images per-channel by subtracting the mean and dividing by the std deviation.\n3.  For Submission, normalize all the submission images per-channel by subtracting the mean and dividing by the std deviation.\n\nAgain, the question is:   Is there enough variation that this is worthwhile?",
    "1215615": "Hi @markalavin Usually normalization leads to faster convergence. So it is always worthwhile to try it out.\nIn my personal experience it does not always deliver on its promises but as said ... try it out.\n\nUse the same normalization for Train, Test and Validation. Good luck.",
    "1218015": "My guess is that it also depends on the architecture. Network like EfficientNet includes batch normalization layers (though not as the first layer), so benefits may not be that large. I also wonder if doing per-channel pre-normalization would introduce additional errors, like a mismatch between values in different channels.",
    "1218116": "I agree with you on the point 'mismatch'. It sounds reasonable.",
    "1218167": "Here's how I did the normalization, and what the results were:\n\nDeviating from the original description, I normalized each image independent of the others, i.e., I did not collect statistics for all the images together.   Rather, for each image, for each channel, I calculated the mean and standard deviation, then subtracted the per-channel means from channel's values and divided it by the per-channel standard deviation and multiplied the result by 255/6 (because I want +/- 3 stddev's to map into [0-255]).\n\nWhen I tried building a model with these normalized images, I found that the out-of-fold validation DICE score was very low compared to what I had seen with the unnormalized images, e.g., 0.6 instead of 0.9.   So, I just gave up on the normalization idea.   But I'd gladly try anything you could suggest."
  },
  "source": "meta"
}