{
  "id": 231164,
  "title": "Normalization values for train dataset",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/231164",
  "author_name": "",
  "post_date": "2021-04-07T08:41:55.070639600Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>A common trick to improve performance of a CNN is to apply channel-wise (i.e. for the red, green, and blue channel) centering and uniform scaling to the training images (similar to what batch-norm does for every batch).</p>\n<p>The formula is simply: <code>img = (img - mean) / std</code></p>\n<p>I have computed the channel-wise mean and std of the training set <a href=\"https://www.kaggle.com/paulgavrikov/normalization-values-for-pp2021\" target=\"_blank\">in this kernel.</a> These values were computed on the full, non-resized training set. The images were transformed to a [0, 1] channel range by division (<code>img / 255</code>). If you resize your images or remove duplicates these values will slightly change but that should be insignificant.</p>\n<p>Here is an example for a PyTorch transform block:</p>\n<pre><code>data_transform = transforms.Compose([\n        ..., # your transformations go here\n        transforms.ToTensor(),\n        transforms.Normalize(mean=[0.48718716, 0.62651698, 0.40807052], \n                             std=[0.17111757, 0.14961184, 0.17837469])\n])\n</code></pre>",
  "messages": [
    {
      "id": "1265831",
      "postDate": "04/07/2021 08:41:55",
      "content": "<p>A common trick to improve performance of a CNN is to apply channel-wise (i.e. for the red, green, and blue channel) centering and uniform scaling to the training images (similar to what batch-norm does for every batch).</p>\n<p>The formula is simply: <code>img = (img - mean) / std</code></p>\n<p>I have computed the channel-wise mean and std of the training set <a href=\"https://www.kaggle.com/paulgavrikov/normalization-values-for-pp2021\" target=\"_blank\">in this kernel.</a> These values were computed on the full, non-resized training set. The images were transformed to a [0, 1] channel range by division (<code>img / 255</code>). If you resize your images or remove duplicates these values will slightly change but that should be insignificant.</p>\n<p>Here is an example for a PyTorch transform block:</p>\n<pre><code>data_transform = transforms.Compose([\n        ..., # your transformations go here\n        transforms.ToTensor(),\n        transforms.Normalize(mean=[0.48718716, 0.62651698, 0.40807052], \n                             std=[0.17111757, 0.14961184, 0.17837469])\n])\n</code></pre>",
      "rawMarkdown": "A common trick to improve performance of a CNN is to apply channel-wise (i.e. for the red, green, and blue channel) centering and uniform scaling to the training images (similar to what batch-norm does for every batch).\n\nThe formula is simply: `img = (img - mean) / std`\n\nI have computed the channel-wise mean and std of the training set [in this kernel.](https://www.kaggle.com/paulgavrikov/normalization-values-for-pp2021) These values were computed on the full, non-resized training set. The images were transformed to a [0, 1] channel range by division (`img / 255`). If you resize your images or remove duplicates these values will slightly change but that should be insignificant.\n\nHere is an example for a PyTorch transform block:\n\n```\ndata_transform = transforms.Compose([\n        ..., # your transformations go here\n        transforms.ToTensor(),\n        transforms.Normalize(mean=[0.48718716, 0.62651698, 0.40807052], \n                             std=[0.17111757, 0.14961184, 0.17837469])\n])\n```",
      "votes": null
    },
    {
      "id": "1266273",
      "postDate": "04/07/2021 16:01:37",
      "content": "<p>If you are going to train model from scratch, i.e., with randomly initialized weights as supposed to per-trained model on some other data, then you do not need to standardize data. This will happen anyways when you train the model since it will apply batch normalization. All what needs to be done is scaling pixel values to 0-1 range by dividing all image tensors by 255.</p>",
      "rawMarkdown": "If you are going to train model from scratch, i.e., with randomly initialized weights as supposed to per-trained model on some other data, then you do not need to standardize data. This will happen anyways when you train the model since it will apply batch normalization. All what needs to be done is scaling pixel values to 0-1 range by dividing all image tensors by 255.",
      "votes": null
    },
    {
      "id": "1267445",
      "postDate": "04/08/2021 14:17:41",
      "content": "<p>I’ve seen a number of implementations that try to squeeze out as much accuracy as possible when training from scratch (eg resnet18 on cifar). And they often of mot always use normalization. I understand your reasoning but perhaps there’s another effect. </p>",
      "rawMarkdown": "I’ve seen a number of implementations that try to squeeze out as much accuracy as possible when training from scratch (eg resnet18 on cifar). And they often of mot always use normalization. I understand your reasoning but perhaps there’s another effect.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1266273,
      "author_name": "datasciencegeek",
      "author_url": "",
      "post_date": "04/07/2021 16:01:37",
      "content": "<p>If you are going to train model from scratch, i.e., with randomly initialized weights as supposed to per-trained model on some other data, then you do not need to standardize data. This will happen anyways when you train the model since it will apply batch normalization. All what needs to be done is scaling pixel values to 0-1 range by dividing all image tensors by 255.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1267445,
          "author_name": "paulgavrikov",
          "author_url": "",
          "post_date": "04/08/2021 14:17:41",
          "content": "<p>I’ve seen a number of implementations that try to squeeze out as much accuracy as possible when training from scratch (eg resnet18 on cifar). And they often of mot always use normalization. I understand your reasoning but perhaps there’s another effect. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1265831": "A common trick to improve performance of a CNN is to apply channel-wise (i.e. for the red, green, and blue channel) centering and uniform scaling to the training images (similar to what batch-norm does for every batch).\n\nThe formula is simply: `img = (img - mean) / std`\n\nI have computed the channel-wise mean and std of the training set [in this kernel.](https://www.kaggle.com/paulgavrikov/normalization-values-for-pp2021) These values were computed on the full, non-resized training set. The images were transformed to a [0, 1] channel range by division (`img / 255`). If you resize your images or remove duplicates these values will slightly change but that should be insignificant.\n\nHere is an example for a PyTorch transform block:\n\n```\ndata_transform = transforms.Compose([\n        ..., # your transformations go here\n        transforms.ToTensor(),\n        transforms.Normalize(mean=[0.48718716, 0.62651698, 0.40807052], \n                             std=[0.17111757, 0.14961184, 0.17837469])\n])\n```",
    "1266273": "If you are going to train model from scratch, i.e., with randomly initialized weights as supposed to per-trained model on some other data, then you do not need to standardize data. This will happen anyways when you train the model since it will apply batch normalization. All what needs to be done is scaling pixel values to 0-1 range by dividing all image tensors by 255.",
    "1267445": "I’ve seen a number of implementations that try to squeeze out as much accuracy as possible when training from scratch (eg resnet18 on cifar). And they often of mot always use normalization. I understand your reasoning but perhaps there’s another effect."
  },
  "source": "meta"
}