{
  "id": 102997,
  "title": "How to normalize properly?",
  "url": "/competitions/recursion-cellular-image-classification/discussion/102997",
  "author_name": "",
  "post_date": "2019-08-06T11:38:55.158997200Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Almost everyone who uses pytorch transforms firstly do ToTensor() and then Normalize(mean=[0.5,...], std =[0.25,...]), but I, after applying just ToTensor() already have analmost zero mean (but not 1 std). And normalizing only making things worse.</p>",
  "messages": [
    {
      "id": "593264",
      "postDate": "08/06/2019 11:38:55",
      "content": "<p>Almost everyone who uses pytorch transforms firstly do ToTensor() and then Normalize(mean=[0.5,...], std =[0.25,...]), but I, after applying just ToTensor() already have analmost zero mean (but not 1 std). And normalizing only making things worse.</p>",
      "rawMarkdown": "Almost everyone who uses pytorch transforms firstly do ToTensor() and then Normalize(mean=[0.5,...], std =[0.25,...]), but I, after applying just ToTensor() already have analmost zero mean (but not 1 std). And normalizing only making things worse.",
      "votes": null
    },
    {
      "id": "593345",
      "postDate": "08/06/2019 13:30:31",
      "content": "<p>What do you mean with 'make things worse'? ToTensor divides the pixel values by 255 so after it is from 0. to 1.\nHowever, you generally want to have your data with mean == 0. and std == 1. \nTherefore, you subtract by the dataset's mean and divide by its standard deviation.</p>",
      "rawMarkdown": "What do you mean with 'make things worse'? ToTensor divides the pixel values by 255 so after it is from 0. to 1.\nHowever, you generally want to have your data with mean == 0. and std == 1. \nTherefore, you subtract by the dataset's mean and divide by its standard deviation.",
      "votes": null
    },
    {
      "id": "593346",
      "postDate": "08/06/2019 13:30:39",
      "content": "<p>Maybe use nn.BatchNorm2d(6) as a first layer of my network?</p>",
      "rawMarkdown": "Maybe use nn.BatchNorm2d(6) as a first layer of my network?",
      "votes": null
    },
    {
      "id": "593351",
      "postDate": "08/06/2019 13:34:39",
      "content": "<p>Because of this my network converges more slowly</p>\n\n<p>After i got values of each channel in [0,1] range, i have a mean of approx. 0 and std of approx 0.05.\nAnd not 0.5 and 0.25.</p>",
      "rawMarkdown": "Because of this my network converges more slowly\n\nAfter i got values of each channel in [0,1] range, i have a mean of approx. 0 and std of approx 0.05.\nAnd not 0.5 and 0.25.",
      "votes": null
    },
    {
      "id": "593360",
      "postDate": "08/06/2019 13:45:27",
      "content": "<p>That would work, but the norm is to get the statistics from the whole data, including test set.</p>",
      "rawMarkdown": "That would work, but the norm is to get the statistics from the whole data, including test set.",
      "votes": null
    },
    {
      "id": "593367",
      "postDate": "08/06/2019 13:58:44",
      "content": "<p>I'd have a look on the public kernels, and maybe just use their dataset/loader and normalisation... I personally prefer starting from a working baseline, and build on top of that, change by change. Making sure you test everything on the way, as deep learning often 'works' even if there is a pretty harsh mistake somewhere. \nAnyway, your inputs to the network should be mean 0 and std 1, which means that they are normally between [and close to] -1 and 1.</p>",
      "rawMarkdown": "I'd have a look on the public kernels, and maybe just use their dataset/loader and normalisation... I personally prefer starting from a working baseline, and build on top of that, change by change. Making sure you test everything on the way, as deep learning often 'works' even if there is a pretty harsh mistake somewhere. \nAnyway, your inputs to the network should be mean 0 and std 1, which means that they are normally between [and close to] -1 and 1.",
      "votes": null
    },
    {
      "id": "610991",
      "postDate": "08/29/2019 03:50:22",
      "content": "<p>I'm having some issues with normalizing the images. I tried different methods but my models are always overfitting (split by cell line).</p>\n\n<p>The following images refer to the same index of the DataFrame. In the title of each image (the 6 channels), I print the minimum, maximum, mean and standard deviation.</p>\n\n<ul>\n<li>Original images\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2F0058efb425024684f937765e6eb742c2%2Foriginal.png?generation=1567050123853789&amp;alt=media\" alt=\"\"></li>\n</ul>\n\n<p>The images are not in the range [0, 1], the mean is not 0 and the std is not 1. And these values change from image to image.</p>\n\n<ul>\n<li>Using the global statistics from the RXRX scripts with <code>Normalize(mean, std)</code>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2Fec85cc981d2b4327b53f6ee35c81da31%2Fglobal.png?generation=1567050322308170&amp;alt=media\" alt=\"\"></li>\n</ul>\n\n<p>I thought that by removing the mean and dividing by the std, the mean and std would become 0 and 1, respectively. But it's not the case here. Am I missing something? Shouldn't we compute these statistics from the training set and apply them to the validation and test sets?</p>\n\n<ul>\n<li>Dividing the images by 255 only\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2F4785b16041d89b680a669fad566b2638%2Fdiv_255.png?generation=1567050451264890&amp;alt=media\" alt=\"\"></li>\n</ul>\n\n<p>In this case the images are indeed in the range [0, 1], more or less. But the mean and std are not 0 and 1, respectively.</p>\n\n<ul>\n<li>Using a <code>BatchNorm2d</code> layer\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2Fabdf373900dc35084a931a083a7dce87%2Fbn.png?generation=1567050530194310&amp;alt=media\" alt=\"\"></li>\n</ul>\n\n<p>In this case the mean and std are indeed 0 and 1, respectively. But now the range is not in [0, 1]. Moreover, this is done per batch and maybe is not suitable to handle the experimental batch effect.</p>\n\n<p>Any suggestion on which approach might be more beneficial?</p>",
      "rawMarkdown": "I'm having some issues with normalizing the images. I tried different methods but my models are always overfitting (split by cell line).\n\nThe following images refer to the same index of the DataFrame. In the title of each image (the 6 channels), I print the minimum, maximum, mean and standard deviation.\n\n- Original images\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2F0058efb425024684f937765e6eb742c2%2Foriginal.png?generation=1567050123853789&amp;alt=media)\n\nThe images are not in the range [0, 1], the mean is not 0 and the std is not 1. And these values change from image to image.\n\n- Using the global statistics from the RXRX scripts with `Normalize(mean, std)`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2Fec85cc981d2b4327b53f6ee35c81da31%2Fglobal.png?generation=1567050322308170&amp;alt=media)\n\nI thought that by removing the mean and dividing by the std, the mean and std would become 0 and 1, respectively. But it's not the case here. Am I missing something? Shouldn't we compute these statistics from the training set and apply them to the validation and test sets?\n\n- Dividing the images by 255 only\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2F4785b16041d89b680a669fad566b2638%2Fdiv_255.png?generation=1567050451264890&amp;alt=media)\n\nIn this case the images are indeed in the range [0, 1], more or less. But the mean and std are not 0 and 1, respectively.\n\n- Using a `BatchNorm2d` layer\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2Fabdf373900dc35084a931a083a7dce87%2Fbn.png?generation=1567050530194310&amp;alt=media)\n\nIn this case the mean and std are indeed 0 and 1, respectively. But now the range is not in [0, 1]. Moreover, this is done per batch and maybe is not suitable to handle the experimental batch effect.\n\nAny suggestion on which approach might be more beneficial?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 593345,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "08/06/2019 13:30:31",
      "content": "<p>What do you mean with 'make things worse'? ToTensor divides the pixel values by 255 so after it is from 0. to 1.\nHowever, you generally want to have your data with mean == 0. and std == 1. \nTherefore, you subtract by the dataset's mean and divide by its standard deviation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 593351,
          "author_name": "rafailfridman",
          "author_url": "",
          "post_date": "08/06/2019 13:34:39",
          "content": "<p>Because of this my network converges more slowly</p>\n\n<p>After i got values of each channel in [0,1] range, i have a mean of approx. 0 and std of approx 0.05.\nAnd not 0.5 and 0.25.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 593367,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "08/06/2019 13:58:44",
          "content": "<p>I'd have a look on the public kernels, and maybe just use their dataset/loader and normalisation... I personally prefer starting from a working baseline, and build on top of that, change by change. Making sure you test everything on the way, as deep learning often 'works' even if there is a pretty harsh mistake somewhere. \nAnyway, your inputs to the network should be mean 0 and std 1, which means that they are normally between [and close to] -1 and 1.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 593346,
      "author_name": "rafailfridman",
      "author_url": "",
      "post_date": "08/06/2019 13:30:39",
      "content": "<p>Maybe use nn.BatchNorm2d(6) as a first layer of my network?</p>",
      "votes": null,
      "replies": [
        {
          "id": 593360,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "08/06/2019 13:45:27",
          "content": "<p>That would work, but the norm is to get the statistics from the whole data, including test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 610991,
      "author_name": "lorenzofabbri92",
      "author_url": "",
      "post_date": "08/29/2019 03:50:22",
      "content": "<p>I'm having some issues with normalizing the images. I tried different methods but my models are always overfitting (split by cell line).</p>\n\n<p>The following images refer to the same index of the DataFrame. In the title of each image (the 6 channels), I print the minimum, maximum, mean and standard deviation.</p>\n\n<ul>\n<li>Original images\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2F0058efb425024684f937765e6eb742c2%2Foriginal.png?generation=1567050123853789&amp;alt=media\" alt=\"\"></li>\n</ul>\n\n<p>The images are not in the range [0, 1], the mean is not 0 and the std is not 1. And these values change from image to image.</p>\n\n<ul>\n<li>Using the global statistics from the RXRX scripts with <code>Normalize(mean, std)</code>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2Fec85cc981d2b4327b53f6ee35c81da31%2Fglobal.png?generation=1567050322308170&amp;alt=media\" alt=\"\"></li>\n</ul>\n\n<p>I thought that by removing the mean and dividing by the std, the mean and std would become 0 and 1, respectively. But it's not the case here. Am I missing something? Shouldn't we compute these statistics from the training set and apply them to the validation and test sets?</p>\n\n<ul>\n<li>Dividing the images by 255 only\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2F4785b16041d89b680a669fad566b2638%2Fdiv_255.png?generation=1567050451264890&amp;alt=media\" alt=\"\"></li>\n</ul>\n\n<p>In this case the images are indeed in the range [0, 1], more or less. But the mean and std are not 0 and 1, respectively.</p>\n\n<ul>\n<li>Using a <code>BatchNorm2d</code> layer\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2Fabdf373900dc35084a931a083a7dce87%2Fbn.png?generation=1567050530194310&amp;alt=media\" alt=\"\"></li>\n</ul>\n\n<p>In this case the mean and std are indeed 0 and 1, respectively. But now the range is not in [0, 1]. Moreover, this is done per batch and maybe is not suitable to handle the experimental batch effect.</p>\n\n<p>Any suggestion on which approach might be more beneficial?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "593264": "Almost everyone who uses pytorch transforms firstly do ToTensor() and then Normalize(mean=[0.5,...], std =[0.25,...]), but I, after applying just ToTensor() already have analmost zero mean (but not 1 std). And normalizing only making things worse.",
    "593345": "What do you mean with 'make things worse'? ToTensor divides the pixel values by 255 so after it is from 0. to 1.\nHowever, you generally want to have your data with mean == 0. and std == 1. \nTherefore, you subtract by the dataset's mean and divide by its standard deviation.",
    "593346": "Maybe use nn.BatchNorm2d(6) as a first layer of my network?",
    "593351": "Because of this my network converges more slowly\n\nAfter i got values of each channel in [0,1] range, i have a mean of approx. 0 and std of approx 0.05.\nAnd not 0.5 and 0.25.",
    "593360": "That would work, but the norm is to get the statistics from the whole data, including test set.",
    "593367": "I'd have a look on the public kernels, and maybe just use their dataset/loader and normalisation... I personally prefer starting from a working baseline, and build on top of that, change by change. Making sure you test everything on the way, as deep learning often 'works' even if there is a pretty harsh mistake somewhere. \nAnyway, your inputs to the network should be mean 0 and std 1, which means that they are normally between [and close to] -1 and 1.",
    "610991": "I'm having some issues with normalizing the images. I tried different methods but my models are always overfitting (split by cell line).\n\nThe following images refer to the same index of the DataFrame. In the title of each image (the 6 channels), I print the minimum, maximum, mean and standard deviation.\n\n- Original images\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2F0058efb425024684f937765e6eb742c2%2Foriginal.png?generation=1567050123853789&amp;alt=media)\n\nThe images are not in the range [0, 1], the mean is not 0 and the std is not 1. And these values change from image to image.\n\n- Using the global statistics from the RXRX scripts with `Normalize(mean, std)`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2Fec85cc981d2b4327b53f6ee35c81da31%2Fglobal.png?generation=1567050322308170&amp;alt=media)\n\nI thought that by removing the mean and dividing by the std, the mean and std would become 0 and 1, respectively. But it's not the case here. Am I missing something? Shouldn't we compute these statistics from the training set and apply them to the validation and test sets?\n\n- Dividing the images by 255 only\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2F4785b16041d89b680a669fad566b2638%2Fdiv_255.png?generation=1567050451264890&amp;alt=media)\n\nIn this case the images are indeed in the range [0, 1], more or less. But the mean and std are not 0 and 1, respectively.\n\n- Using a `BatchNorm2d` layer\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2Fabdf373900dc35084a931a083a7dce87%2Fbn.png?generation=1567050530194310&amp;alt=media)\n\nIn this case the mean and std are indeed 0 and 1, respectively. But now the range is not in [0, 1]. Moreover, this is done per batch and maybe is not suitable to handle the experimental batch effect.\n\nAny suggestion on which approach might be more beneficial?"
  },
  "source": "meta"
}