{
  "id": 130640,
  "title": "Image Normalization",
  "url": "/competitions/deepfake-detection-challenge/discussion/130640",
  "author_name": "",
  "post_date": "2020-02-15T13:38:46.176999300Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi, I am new to image processing. So I wonder what's the purpose of the normalization I have seen in the @humananalog kernel? Is this a common practice and why the 3 RGB channels have got different parameters? </p>\n\n<p>from torchvision.transforms import Normalize</p>\n\n<p>mean = [0.485, 0.456, 0.406]\nstd = [0.229, 0.224, 0.225]\nnormalize_transform = Normalize(mean, std)</p>",
  "messages": [
    {
      "id": "746740",
      "postDate": "02/15/2020 13:38:46",
      "content": "<p>Hi, I am new to image processing. So I wonder what's the purpose of the normalization I have seen in the @humananalog kernel? Is this a common practice and why the 3 RGB channels have got different parameters? </p>\n\n<p>from torchvision.transforms import Normalize</p>\n\n<p>mean = [0.485, 0.456, 0.406]\nstd = [0.229, 0.224, 0.225]\nnormalize_transform = Normalize(mean, std)</p>",
      "rawMarkdown": "Hi, I am new to image processing. So I wonder what's the purpose of the normalization I have seen in the @humananalog kernel? Is this a common practice and why the 3 RGB channels have got different parameters? \n\nfrom torchvision.transforms import Normalize\n\nmean = [0.485, 0.456, 0.406]\nstd = [0.229, 0.224, 0.225]\nnormalize_transform = Normalize(mean, std)",
      "votes": null
    },
    {
      "id": "746761",
      "postDate": "02/15/2020 14:04:41",
      "content": "<p>Neural networks tend to do better when their inputs have a mean of 0 and a standard deviation of 1. For this reason, the 'pretrained' networks have been trained with images processed this way. If you therefore want to use those pretrained networks as a starting point, you need to make your images are transformed like that, too. Those values are the mean and standard deviations per colour channel of the ImageNet dataset, which models are classically pretrained on.</p>\n\n<p>There are quite a few reasons we do that when training a network (ie why they're pretrained that way), but they include that when the network is randomly initialised, they're initialised in a pattern where the magnitude of inputs are expected to be in that range. Also, we often use weight decay where very strong weights are penalised, and again these numbers are in a chosen to work well inputs in the standard range.</p>",
      "rawMarkdown": "Neural networks tend to do better when their inputs have a mean of 0 and a standard deviation of 1. For this reason, the 'pretrained' networks have been trained with images processed this way. If you therefore want to use those pretrained networks as a starting point, you need to make your images are transformed like that, too. Those values are the mean and standard deviations per colour channel of the ImageNet dataset, which models are classically pretrained on.\n\nThere are quite a few reasons we do that when training a network (ie why they're pretrained that way), but they include that when the network is randomly initialised, they're initialised in a pattern where the magnitude of inputs are expected to be in that range. Also, we often use weight decay where very strong weights are penalised, and again these numbers are in a chosen to work well inputs in the standard range.",
      "votes": null
    },
    {
      "id": "746788",
      "postDate": "02/15/2020 14:51:28",
      "content": "<p>Thanks James, makes sense</p>",
      "rawMarkdown": "Thanks James, makes sense",
      "votes": null
    },
    {
      "id": "747048",
      "postDate": "02/15/2020 22:06:56",
      "content": "<p>I agree with your comments. I think, though that one must compute the mean and std on one's own dataset rather than using the hardwired numbers above.</p>",
      "rawMarkdown": "I agree with your comments. I think, though that one must compute the mean and std on one's own dataset rather than using the hardwired numbers above.",
      "votes": null
    },
    {
      "id": "747051",
      "postDate": "02/15/2020 22:16:54",
      "content": "<p>It depends on what you want to do. If you want to train the network from scratch, then I agree that is sensible.\nHowever, if you are using transfer learning and freezing the first layers of the network, then it's probably better to use the normalization parameters of the original dataset.</p>\n\n<p>For example, let's say you want to use transfer learning to refashion a network trained on ImageNet into a network to classify breeds cats. Cats don't come in blue or green, and so these colours channels will be under-represented, whereas browns, oranges and so on do exist. The mean and std of your dataset will therefore be red-biased.</p>\n\n<p>Now neural networks trained on ImageNet are actually very good at classifying animals (especially dogs), and freezing an ImageNet trained network and so using the features coming out of the later layers to create a cat classifier will probably work really well.</p>\n\n<p>However, if you renormalize your images using your cat dataset's mean/std then what the pre-trained network has done to recognise brown will no longer work -  a brown image from your dataset will look grey to the network, because the channels become rebalanced.</p>\n\n<p>What this means is the images you feed in will now go down different parts of the network than they would have done if you'd have used Imagenet normalisation parameters.</p>\n\n<p>It's a contrived example, I know, but I hope it illustrates the point.</p>\n\n<p>Jeremy Howard makes this point in his fastai course (probably better than I've done, but I can't remember how he explained it...)</p>",
      "rawMarkdown": "It depends on what you want to do. If you want to train the network from scratch, then I agree that is sensible.\nHowever, if you are using transfer learning and freezing the first layers of the network, then it's probably better to use the normalization parameters of the original dataset.\n\nFor example, let's say you want to use transfer learning to refashion a network trained on ImageNet into a network to classify breeds cats. Cats don't come in blue or green, and so these colours channels will be under-represented, whereas browns, oranges and so on do exist. The mean and std of your dataset will therefore be red-biased.\n\nNow neural networks trained on ImageNet are actually very good at classifying animals (especially dogs), and freezing an ImageNet trained network and so using the features coming out of the later layers to create a cat classifier will probably work really well.\n\nHowever, if you renormalize your images using your cat dataset's mean/std then what the pre-trained network has done to recognise brown will no longer work -  a brown image from your dataset will look grey to the network, because the channels become rebalanced.\n\nWhat this means is the images you feed in will now go down different parts of the network than they would have done if you'd have used Imagenet normalisation parameters.\n\nIt's a contrived example, I know, but I hope it illustrates the point.\n\nJeremy Howard makes this point in his fastai course (probably better than I've done, but I can't remember how he explained it...)",
      "votes": null
    },
    {
      "id": "747127",
      "postDate": "02/16/2020 01:48:03",
      "content": "<p>\"Normalize\" in the context of neural networks means \"subtract the mean and then divide by standard deviation\" where the mean and std are those of the data set you are working with. That is exactly what the code in, say, pytorch \"normalize\" does when you call it.When you do that, you end up with a mean/std that are 0 and 1 respectively.\n That is what the imagenet people did with their data, so the data they feed into their network (the network that we use for pretraining) has mean/std = 0, 1.\n To use their NN weights, we need to make our data have the same mean and std  (0,1) and the only way to do that is to normalize with our own statistics.</p>\n\n<p>If we would use their numbers in normalizing, we would end up with some wierd combination of our numbers and theirs.</p>\n\n<p>Note that the individual image  std's can have a LOT of variance - it is just the stats over the entire data set we work with.</p>\n\n<p>That's my take on it anyway.</p>",
      "rawMarkdown": "\"Normalize\" in the context of neural networks means \"subtract the mean and then divide by standard deviation\" where the mean and std are those of the data set you are working with. That is exactly what the code in, say, pytorch \"normalize\" does when you call it.When you do that, you end up with a mean/std that are 0 and 1 respectively.\n That is what the imagenet people did with their data, so the data they feed into their network (the network that we use for pretraining) has mean/std = 0, 1.\n To use their NN weights, we need to make our data have the same mean and std  (0,1) and the only way to do that is to normalize with our own statistics.\n\nIf we would use their numbers in normalizing, we would end up with some wierd combination of our numbers and theirs.\n\n  Note that the individual image  std's can have a LOT of variance - it is just the stats over the entire data set we work with.\n\n  That's my take on it anyway.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 746761,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "02/15/2020 14:04:41",
      "content": "<p>Neural networks tend to do better when their inputs have a mean of 0 and a standard deviation of 1. For this reason, the 'pretrained' networks have been trained with images processed this way. If you therefore want to use those pretrained networks as a starting point, you need to make your images are transformed like that, too. Those values are the mean and standard deviations per colour channel of the ImageNet dataset, which models are classically pretrained on.</p>\n\n<p>There are quite a few reasons we do that when training a network (ie why they're pretrained that way), but they include that when the network is randomly initialised, they're initialised in a pattern where the magnitude of inputs are expected to be in that range. Also, we often use weight decay where very strong weights are penalised, and again these numbers are in a chosen to work well inputs in the standard range.</p>",
      "votes": null,
      "replies": [
        {
          "id": 747048,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "02/15/2020 22:06:56",
          "content": "<p>I agree with your comments. I think, though that one must compute the mean and std on one's own dataset rather than using the hardwired numbers above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747051,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "02/15/2020 22:16:54",
          "content": "<p>It depends on what you want to do. If you want to train the network from scratch, then I agree that is sensible.\nHowever, if you are using transfer learning and freezing the first layers of the network, then it's probably better to use the normalization parameters of the original dataset.</p>\n\n<p>For example, let's say you want to use transfer learning to refashion a network trained on ImageNet into a network to classify breeds cats. Cats don't come in blue or green, and so these colours channels will be under-represented, whereas browns, oranges and so on do exist. The mean and std of your dataset will therefore be red-biased.</p>\n\n<p>Now neural networks trained on ImageNet are actually very good at classifying animals (especially dogs), and freezing an ImageNet trained network and so using the features coming out of the later layers to create a cat classifier will probably work really well.</p>\n\n<p>However, if you renormalize your images using your cat dataset's mean/std then what the pre-trained network has done to recognise brown will no longer work -  a brown image from your dataset will look grey to the network, because the channels become rebalanced.</p>\n\n<p>What this means is the images you feed in will now go down different parts of the network than they would have done if you'd have used Imagenet normalisation parameters.</p>\n\n<p>It's a contrived example, I know, but I hope it illustrates the point.</p>\n\n<p>Jeremy Howard makes this point in his fastai course (probably better than I've done, but I can't remember how he explained it...)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747127,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "02/16/2020 01:48:03",
          "content": "<p>\"Normalize\" in the context of neural networks means \"subtract the mean and then divide by standard deviation\" where the mean and std are those of the data set you are working with. That is exactly what the code in, say, pytorch \"normalize\" does when you call it.When you do that, you end up with a mean/std that are 0 and 1 respectively.\n That is what the imagenet people did with their data, so the data they feed into their network (the network that we use for pretraining) has mean/std = 0, 1.\n To use their NN weights, we need to make our data have the same mean and std  (0,1) and the only way to do that is to normalize with our own statistics.</p>\n\n<p>If we would use their numbers in normalizing, we would end up with some wierd combination of our numbers and theirs.</p>\n\n<p>Note that the individual image  std's can have a LOT of variance - it is just the stats over the entire data set we work with.</p>\n\n<p>That's my take on it anyway.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 746788,
      "author_name": "antonio1979",
      "author_url": "",
      "post_date": "02/15/2020 14:51:28",
      "content": "<p>Thanks James, makes sense</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "746740": "Hi, I am new to image processing. So I wonder what's the purpose of the normalization I have seen in the @humananalog kernel? Is this a common practice and why the 3 RGB channels have got different parameters? \n\nfrom torchvision.transforms import Normalize\n\nmean = [0.485, 0.456, 0.406]\nstd = [0.229, 0.224, 0.225]\nnormalize_transform = Normalize(mean, std)",
    "746761": "Neural networks tend to do better when their inputs have a mean of 0 and a standard deviation of 1. For this reason, the 'pretrained' networks have been trained with images processed this way. If you therefore want to use those pretrained networks as a starting point, you need to make your images are transformed like that, too. Those values are the mean and standard deviations per colour channel of the ImageNet dataset, which models are classically pretrained on.\n\nThere are quite a few reasons we do that when training a network (ie why they're pretrained that way), but they include that when the network is randomly initialised, they're initialised in a pattern where the magnitude of inputs are expected to be in that range. Also, we often use weight decay where very strong weights are penalised, and again these numbers are in a chosen to work well inputs in the standard range.",
    "746788": "Thanks James, makes sense",
    "747048": "I agree with your comments. I think, though that one must compute the mean and std on one's own dataset rather than using the hardwired numbers above.",
    "747051": "It depends on what you want to do. If you want to train the network from scratch, then I agree that is sensible.\nHowever, if you are using transfer learning and freezing the first layers of the network, then it's probably better to use the normalization parameters of the original dataset.\n\nFor example, let's say you want to use transfer learning to refashion a network trained on ImageNet into a network to classify breeds cats. Cats don't come in blue or green, and so these colours channels will be under-represented, whereas browns, oranges and so on do exist. The mean and std of your dataset will therefore be red-biased.\n\nNow neural networks trained on ImageNet are actually very good at classifying animals (especially dogs), and freezing an ImageNet trained network and so using the features coming out of the later layers to create a cat classifier will probably work really well.\n\nHowever, if you renormalize your images using your cat dataset's mean/std then what the pre-trained network has done to recognise brown will no longer work -  a brown image from your dataset will look grey to the network, because the channels become rebalanced.\n\nWhat this means is the images you feed in will now go down different parts of the network than they would have done if you'd have used Imagenet normalisation parameters.\n\nIt's a contrived example, I know, but I hope it illustrates the point.\n\nJeremy Howard makes this point in his fastai course (probably better than I've done, but I can't remember how he explained it...)",
    "747127": "\"Normalize\" in the context of neural networks means \"subtract the mean and then divide by standard deviation\" where the mean and std are those of the data set you are working with. That is exactly what the code in, say, pytorch \"normalize\" does when you call it.When you do that, you end up with a mean/std that are 0 and 1 respectively.\n That is what the imagenet people did with their data, so the data they feed into their network (the network that we use for pretraining) has mean/std = 0, 1.\n To use their NN weights, we need to make our data have the same mean and std  (0,1) and the only way to do that is to normalize with our own statistics.\n\nIf we would use their numbers in normalizing, we would end up with some wierd combination of our numbers and theirs.\n\n  Note that the individual image  std's can have a LOT of variance - it is just the stats over the entire data set we work with.\n\n  That's my take on it anyway."
  },
  "source": "meta"
}