{
  "id": 211039,
  "title": "Normalization Constants",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/211039",
  "author_name": "",
  "post_date": "2021-01-13T11:43:21.834698800Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>If you want to normalize the one input channel of the X-ray images, I have calculated the mean and standard deviation constants over the entire dataset :)</p>\n<p>mean=[0.4824], std=[0.234]</p>\n<p>NOTE: My initial reported STD value was incorrect! Corrected as of 21 February 2021.</p>",
  "messages": [
    {
      "id": "1151521",
      "postDate": "01/13/2021 11:43:21",
      "content": "<p>If you want to normalize the one input channel of the X-ray images, I have calculated the mean and standard deviation constants over the entire dataset :)</p>\n<p>mean=[0.4824], std=[0.234]</p>\n<p>NOTE: My initial reported STD value was incorrect! Corrected as of 21 February 2021.</p>",
      "rawMarkdown": "If you want to normalize the one input channel of the X-ray images, I have calculated the mean and standard deviation constants over the entire dataset :)\n\nmean=[0.4824], std=[0.234]\n\nNOTE: My initial reported STD value was incorrect! Corrected as of 21 February 2021.",
      "votes": null
    },
    {
      "id": "1212685",
      "postDate": "02/21/2021 13:34:44",
      "content": "<p>Can I double check whether those numbers are the ones one should use? The <code>std</code> you quote seems more like the variance in the means?! Should that not be more like 0.22?</p>\n<p>Or am I getting this wrong / looking at the wrong thing that one should standardize to?</p>\n<pre><code>from PIL import Image\nfrom PIL import ImageFile\nfrom tqdm import tqdm\nimport numpy as np\nimport pandas as pd\nimport fastcore\nfrom fastcore.parallel import parallel\n\ntrain_df = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/train.csv')\nimages = [ f'../input/ranzcr-clip-catheter-line-classification/train/{img}.jpg' for img in train_df['StudyInstanceUID']]\n\ndef get_msd(imgname):\n    image = Image.open(imgname)\n    return [np.mean(np.array(image)/255.0), np.std(np.array(image)/255.0)]\n\n# tmp = [get_msd(imgname) for imgname in tqdm(images)] # This is a bit slow and takes almost an hour\ntmp = parallel(get_msd, [imgname for imgname in images], n_workers=4, progress=True ) # Parallel processing gets it to about 20 min\n\ninfodf = pd.DataFrame({'mean': [tmpi[0] for tmpi in tmp],\n                       'std': [tmpi[1] for tmpi in tmp]})\n\ninfodf['mean'].mean() # 0.4821912944015841\ninfodf['mean'].var() # 0.005452590552384386\ninfodf['std'].mean() # 0.21995657835377008\n</code></pre>",
      "rawMarkdown": "Can I double check whether those numbers are the ones one should use? The `std` you quote seems more like the variance in the means?! Should that not be more like 0.22?\n\nOr am I getting this wrong / looking at the wrong thing that one should standardize to?\n\n```\nfrom PIL import Image\nfrom PIL import ImageFile\nfrom tqdm import tqdm\nimport numpy as np\nimport pandas as pd\nimport fastcore\nfrom fastcore.parallel import parallel\n\ntrain_df = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/train.csv')\nimages = [ f'../input/ranzcr-clip-catheter-line-classification/train/{img}.jpg' for img in train_df['StudyInstanceUID']]\n\ndef get_msd(imgname):\n    image = Image.open(imgname)\n    return [np.mean(np.array(image)/255.0), np.std(np.array(image)/255.0)]\n\n# tmp = [get_msd(imgname) for imgname in tqdm(images)] # This is a bit slow and takes almost an hour\ntmp = parallel(get_msd, [imgname for imgname in images], n_workers=4, progress=True ) # Parallel processing gets it to about 20 min\n\ninfodf = pd.DataFrame({'mean': [tmpi[0] for tmpi in tmp],\n                       'std': [tmpi[1] for tmpi in tmp]})\n\ninfodf['mean'].mean() # 0.4821912944015841\ninfodf['mean'].var() # 0.005452590552384386\ninfodf['std'].mean() # 0.21995657835377008\n```",
      "votes": null
    },
    {
      "id": "1212758",
      "postDate": "02/21/2021 14:55:48",
      "content": "<p>Hmm, I think you are right. I might have made a mistake. I will quickly double check this and report back.</p>",
      "rawMarkdown": "Hmm, I think you are right. I might have made a mistake. I will quickly double check this and report back.",
      "votes": null
    },
    {
      "id": "1212765",
      "postDate": "02/21/2021 15:10:43",
      "content": "<p>You were right! Thank you so much for correcting me here. I had a bug in my code and accidentally reported the wrong value. I recalculated everything and the updated value is now corrected in the post above. 0.234 is the correct value!</p>",
      "rawMarkdown": "You were right! Thank you so much for correcting me here. I had a bug in my code and accidentally reported the wrong value. I recalculated everything and the updated value is now corrected in the post above. 0.234 is the correct value!",
      "votes": null
    },
    {
      "id": "1212833",
      "postDate": "02/21/2021 16:44:58",
      "content": "<p>Good to have confirmation, thanks. Would have been a mess, if I got that wrong. 👍</p>",
      "rawMarkdown": "Good to have confirmation, thanks. Would have been a mess, if I got that wrong. 👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1212685,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "02/21/2021 13:34:44",
      "content": "<p>Can I double check whether those numbers are the ones one should use? The <code>std</code> you quote seems more like the variance in the means?! Should that not be more like 0.22?</p>\n<p>Or am I getting this wrong / looking at the wrong thing that one should standardize to?</p>\n<pre><code>from PIL import Image\nfrom PIL import ImageFile\nfrom tqdm import tqdm\nimport numpy as np\nimport pandas as pd\nimport fastcore\nfrom fastcore.parallel import parallel\n\ntrain_df = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/train.csv')\nimages = [ f'../input/ranzcr-clip-catheter-line-classification/train/{img}.jpg' for img in train_df['StudyInstanceUID']]\n\ndef get_msd(imgname):\n    image = Image.open(imgname)\n    return [np.mean(np.array(image)/255.0), np.std(np.array(image)/255.0)]\n\n# tmp = [get_msd(imgname) for imgname in tqdm(images)] # This is a bit slow and takes almost an hour\ntmp = parallel(get_msd, [imgname for imgname in images], n_workers=4, progress=True ) # Parallel processing gets it to about 20 min\n\ninfodf = pd.DataFrame({'mean': [tmpi[0] for tmpi in tmp],\n                       'std': [tmpi[1] for tmpi in tmp]})\n\ninfodf['mean'].mean() # 0.4821912944015841\ninfodf['mean'].var() # 0.005452590552384386\ninfodf['std'].mean() # 0.21995657835377008\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1212758,
          "author_name": "raivokoot",
          "author_url": "",
          "post_date": "02/21/2021 14:55:48",
          "content": "<p>Hmm, I think you are right. I might have made a mistake. I will quickly double check this and report back.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1212765,
          "author_name": "raivokoot",
          "author_url": "",
          "post_date": "02/21/2021 15:10:43",
          "content": "<p>You were right! Thank you so much for correcting me here. I had a bug in my code and accidentally reported the wrong value. I recalculated everything and the updated value is now corrected in the post above. 0.234 is the correct value!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1212833,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "02/21/2021 16:44:58",
          "content": "<p>Good to have confirmation, thanks. Would have been a mess, if I got that wrong. 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1151521": "If you want to normalize the one input channel of the X-ray images, I have calculated the mean and standard deviation constants over the entire dataset :)\n\nmean=[0.4824], std=[0.234]\n\nNOTE: My initial reported STD value was incorrect! Corrected as of 21 February 2021.",
    "1212685": "Can I double check whether those numbers are the ones one should use? The `std` you quote seems more like the variance in the means?! Should that not be more like 0.22?\n\nOr am I getting this wrong / looking at the wrong thing that one should standardize to?\n\n```\nfrom PIL import Image\nfrom PIL import ImageFile\nfrom tqdm import tqdm\nimport numpy as np\nimport pandas as pd\nimport fastcore\nfrom fastcore.parallel import parallel\n\ntrain_df = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/train.csv')\nimages = [ f'../input/ranzcr-clip-catheter-line-classification/train/{img}.jpg' for img in train_df['StudyInstanceUID']]\n\ndef get_msd(imgname):\n    image = Image.open(imgname)\n    return [np.mean(np.array(image)/255.0), np.std(np.array(image)/255.0)]\n\n# tmp = [get_msd(imgname) for imgname in tqdm(images)] # This is a bit slow and takes almost an hour\ntmp = parallel(get_msd, [imgname for imgname in images], n_workers=4, progress=True ) # Parallel processing gets it to about 20 min\n\ninfodf = pd.DataFrame({'mean': [tmpi[0] for tmpi in tmp],\n                       'std': [tmpi[1] for tmpi in tmp]})\n\ninfodf['mean'].mean() # 0.4821912944015841\ninfodf['mean'].var() # 0.005452590552384386\ninfodf['std'].mean() # 0.21995657835377008\n```",
    "1212758": "Hmm, I think you are right. I might have made a mistake. I will quickly double check this and report back.",
    "1212765": "You were right! Thank you so much for correcting me here. I had a bug in my code and accidentally reported the wrong value. I recalculated everything and the updated value is now corrected in the post above. 0.234 is the correct value!",
    "1212833": "Good to have confirmation, thanks. Would have been a mess, if I got that wrong. 👍"
  },
  "source": "meta"
}