{
  "id": 232025,
  "title": "Sharing Mean and Standard deviation of images",
  "url": "/competitions/bms-molecular-translation/discussion/232025",
  "author_name": "Gabriel Lindenmaier",
  "post_date": "2021-04-11T19:17:23.905000",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I have seen that some public notebooks use the mean &amp; std from ImageNet.<br>\nI have computed the mean and standard deviation of the competitions training images.<br>\nThey were resized to 192x192, but I doubt that it makes a significant difference.</p>\n<p><strong>Mean:</strong> 251.6988<br>\n**Std:  **  22.8692</p>\n<p>For image loading libs which load them with values between [0, 1] you should divide the values by 255.<br>\nThe values are the same for all channels. As you can see, the majority of the images is white (as it should be)</p>",
  "messages": [
    {
      "id": 1270587,
      "postDate": "2021-04-11T19:17:23.907Z",
      "content": "<p>I have seen that some public notebooks use the mean &amp; std from ImageNet.<br>\nI have computed the mean and standard deviation of the competitions training images.<br>\nThey were resized to 192x192, but I doubt that it makes a significant difference.</p>\n<p><strong>Mean:</strong> 251.6988<br>\n**Std:  **  22.8692</p>\n<p>For image loading libs which load them with values between [0, 1] you should divide the values by 255.<br>\nThe values are the same for all channels. As you can see, the majority of the images is white (as it should be)</p>",
      "rawMarkdown": "I have seen that some public notebooks use the mean & std from ImageNet.\nI have computed the mean and standard deviation of the competitions training images.\nThey were resized to 192x192, but I doubt that it makes a significant difference.\n\n**Mean:** 251.6988\n**Std:  **  22.8692\n\nFor image loading libs which load them with values between [0, 1] you should divide the values by 255.\nThe values are the same for all channels. As you can see, the majority of the images is white (as it should be)",
      "votes": 6
    },
    {
      "id": 1279678,
      "postDate": "2021-04-21T06:35:07.840Z",
      "content": "<p>Can you describe the effects you observed using your calculated values? any metrics to share?</p>",
      "rawMarkdown": "Can you describe the effects you observed using your calculated values? any metrics to share?",
      "replies": [
        {
          "id": 1280112,
          "postDate": "2021-04-21T15:18:23.150Z",
          "content": "<p>Validation performance became stable and stopped fluctuating all over the place. For some runs validation set would have loss of 0.2 after one epoch, other runs 0.7…</p>\n<p>Here is the direct comparison between the last run with wrong std &amp; mean and the fixed version after one epoch:<br>\nWrong:<br>\ntrain - averaged over epoch - top1-accuracy: 0.863 | loss: 0.360<br>\nvalidation -                               top1-accuracy: 0.756 | loss: 0.819</p>\n<p>Fixed:<br>\ntrain - averaged over epoch - top1-accuracy: 0.861 | loss: 0.367<br>\nvalidation -                               top1-accuracy: 0.941 | loss: 0.151</p>\n<p>Accuracy uses masked padding indices, so it is the true top1 accuracy.<br>\nAs you can see only validation performance changed effectively.</p>\n<p>The important part is that your input has mean of 0 and std/variance of 1 on average. That is what you achieve with the correct numbers.</p>\n<p>TLDR: Stabilized validation set performance.</p>",
          "rawMarkdown": "Validation performance became stable and stopped fluctuating all over the place. For some runs validation set would have loss of 0.2 after one epoch, other runs 0.7...\n\nHere is the direct comparison between the last run with wrong std & mean and the fixed version after one epoch:\nWrong:\ntrain - averaged over epoch - top1-accuracy: 0.863 | loss: 0.360\nvalidation -                               top1-accuracy: 0.756 | loss: 0.819\n\nFixed:\ntrain - averaged over epoch - top1-accuracy: 0.861 | loss: 0.367\nvalidation -                               top1-accuracy: 0.941 | loss: 0.151\n\nAccuracy uses masked padding indices, so it is the true top1 accuracy.\nAs you can see only validation performance changed effectively.\n\nThe important part is that your input has mean of 0 and std/variance of 1 on average. That is what you achieve with the correct numbers.\n\nTLDR: Stabilized validation set performance.",
          "votes": 1
        },
        {
          "id": 1280564,
          "postDate": "2021-04-22T05:53:41.930Z",
          "content": "<p>Good feedback, thanks.</p>\n<p>By changing the mean and std values you mean replace the current ImageNet means (0.485,0.456,0.406) and std (0.229,0.224,0.225) with your values divided 255?</p>\n<p>I'm wondering what would happen if you switched the values midway, after training already a certain number of epochs too. You could argue that after a certain number of epochs, the problem is minimized however.</p>",
          "rawMarkdown": "Good feedback, thanks.\n\nBy changing the mean and std values you mean replace the current ImageNet means (0.485,0.456,0.406) and std (0.229,0.224,0.225) with your values divided 255?\n\nI'm wondering what would happen if you switched the values midway, after training already a certain number of epochs too. You could argue that after a certain number of epochs, the problem is minimized however.",
          "votes": 1
        },
        {
          "id": 1280950,
          "postDate": "2021-04-22T14:03:14.700Z",
          "content": "<p>If the library you use to load the images loads them with float values in [0, 1] interval, then divide by 255. If you use torchvision which keeps them in [0, 255] then not. Check beforehand.</p>\n<p>Switching halfway would change your input features considerably, there is no benefit. It should perform slightly worse in the end then using the same one from the start.</p>\n<p>You do the normalizing to bring the values to a smaller, manageable scale and also very important: a stable distribution and especially variance of 1.<br>\nLook into neural network introductions which cover the math part for more insight.</p>",
          "rawMarkdown": "If the library you use to load the images loads them with float values in [0, 1] interval, then divide by 255. If you use torchvision which keeps them in [0, 255] then not. Check beforehand.\n\nSwitching halfway would change your input features considerably, there is no benefit. It should perform slightly worse in the end then using the same one from the start.\n\nYou do the normalizing to bring the values to a smaller, manageable scale and also very important: a stable distribution and especially variance of 1.\nLook into neural network introductions which cover the math part for more insight.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1273259,
      "postDate": "2021-04-14T07:49:54.823Z",
      "content": "<p>Can I know why 255?</p>",
      "rawMarkdown": "Can I know why 255?",
      "replies": [
        {
          "id": 1275985,
          "postDate": "2021-04-17T00:22:30.477Z",
          "content": "<p>Because the colour of the images is in case of RGB (or normal Grayscale) encoding represented by a [0, 255] interval. And that's the basis I used for calculating mean &amp; std.</p>\n<p>If you now load the images with values between [0, 1] you need to normalize the above given mean &amp; std.  So 255/255 = 1. This way we achieve the [0, 1] interval for the mean &amp; std which we need to match the [0, 1] interval of the images.</p>",
          "rawMarkdown": "Because the colour of the images is in case of RGB (or normal Grayscale) encoding represented by a [0, 255] interval. And that's the basis I used for calculating mean & std.\n\nIf you now load the images with values between [0, 1] you need to normalize the above given mean & std.  So 255/255 = 1. This way we achieve the [0, 1] interval for the mean & std which we need to match the [0, 1] interval of the images.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1272277,
      "postDate": "2021-04-13T11:30:17.173Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1272272,
      "postDate": "2021-04-13T11:25:04.397Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1279678,
      "author_name": "Yijie Xu",
      "author_url": "",
      "post_date": "2021-04-21T06:35:07.840000",
      "content": "<p>Can you describe the effects you observed using your calculated values? any metrics to share?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1280112,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-21T15:18:23.150000",
          "content": "<p>Validation performance became stable and stopped fluctuating all over the place. For some runs validation set would have loss of 0.2 after one epoch, other runs 0.7…</p>\n<p>Here is the direct comparison between the last run with wrong std &amp; mean and the fixed version after one epoch:<br>\nWrong:<br>\ntrain - averaged over epoch - top1-accuracy: 0.863 | loss: 0.360<br>\nvalidation -                               top1-accuracy: 0.756 | loss: 0.819</p>\n<p>Fixed:<br>\ntrain - averaged over epoch - top1-accuracy: 0.861 | loss: 0.367<br>\nvalidation -                               top1-accuracy: 0.941 | loss: 0.151</p>\n<p>Accuracy uses masked padding indices, so it is the true top1 accuracy.<br>\nAs you can see only validation performance changed effectively.</p>\n<p>The important part is that your input has mean of 0 and std/variance of 1 on average. That is what you achieve with the correct numbers.</p>\n<p>TLDR: Stabilized validation set performance.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1280564,
          "author_name": "Yijie Xu",
          "author_url": "",
          "post_date": "2021-04-22T05:53:41.930000",
          "content": "<p>Good feedback, thanks.</p>\n<p>By changing the mean and std values you mean replace the current ImageNet means (0.485,0.456,0.406) and std (0.229,0.224,0.225) with your values divided 255?</p>\n<p>I'm wondering what would happen if you switched the values midway, after training already a certain number of epochs too. You could argue that after a certain number of epochs, the problem is minimized however.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1280950,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-22T14:03:14.700000",
          "content": "<p>If the library you use to load the images loads them with float values in [0, 1] interval, then divide by 255. If you use torchvision which keeps them in [0, 255] then not. Check beforehand.</p>\n<p>Switching halfway would change your input features considerably, there is no benefit. It should perform slightly worse in the end then using the same one from the start.</p>\n<p>You do the normalizing to bring the values to a smaller, manageable scale and also very important: a stable distribution and especially variance of 1.<br>\nLook into neural network introductions which cover the math part for more insight.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1273259,
      "author_name": "ASHVIN",
      "author_url": "",
      "post_date": "2021-04-14T07:49:54.823000",
      "content": "<p>Can I know why 255?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1275985,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-17T00:22:30.477000",
          "content": "<p>Because the colour of the images is in case of RGB (or normal Grayscale) encoding represented by a [0, 255] interval. And that's the basis I used for calculating mean &amp; std.</p>\n<p>If you now load the images with values between [0, 1] you need to normalize the above given mean &amp; std.  So 255/255 = 1. This way we achieve the [0, 1] interval for the mean &amp; std which we need to match the [0, 1] interval of the images.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1272277,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-13T11:30:17.173000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1272272,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-13T11:25:04.397000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1270587": "I have seen that some public notebooks use the mean & std from ImageNet.\nI have computed the mean and standard deviation of the competitions training images.\nThey were resized to 192x192, but I doubt that it makes a significant difference.\n\n**Mean:** 251.6988\n**Std:  **  22.8692\n\nFor image loading libs which load them with values between [0, 1] you should divide the values by 255.\nThe values are the same for all channels. As you can see, the majority of the images is white (as it should be)",
    "1279678": "Can you describe the effects you observed using your calculated values? any metrics to share?",
    "1273259": "Can I know why 255?",
    "1272277": "",
    "1272272": ""
  }
}