{
  "id": 107747,
  "title": "do we need to use the pixel_stats.csv?",
  "url": "/competitions/recursion-cellular-image-classification/discussion/107747",
  "author_name": "",
  "post_date": "2019-09-06T11:20:06.528079800Z",
  "votes": 3,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Do we need to normalize EACH of our image with the specific mean and std from the pixel_stats.csv?</p>\n\n<p>Usually I will normalize the whole dataset by pass the same mean and std by torchvision.transforms in pytorch. I am newbie and first time for me to come into this situation.</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "619623",
      "postDate": "09/06/2019 11:20:06",
      "content": "<p>Do we need to normalize EACH of our image with the specific mean and std from the pixel_stats.csv?</p>\n\n<p>Usually I will normalize the whole dataset by pass the same mean and std by torchvision.transforms in pytorch. I am newbie and first time for me to come into this situation.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Do we need to normalize EACH of our image with the specific mean and std from the pixel_stats.csv?\n\nUsually I will normalize the whole dataset by pass the same mean and std by torchvision.transforms in pytorch. I am newbie and first time for me to come into this situation.\n\nThanks",
      "votes": null
    },
    {
      "id": "619890",
      "postDate": "09/06/2019 17:57:44",
      "content": "<p>I think so. One of the themes of this competition is to predict robustly for data from different experiments (batch effect). I am not using pixel_stats.csv but computing a normalization factor for each image each channel. Visual impression from RGB visualization seems that the relative strength between channels also depends on the experiment.</p>",
      "rawMarkdown": "I think so. One of the themes of this competition is to predict robustly for data from different experiments (batch effect). I am not using pixel_stats.csv but computing a normalization factor for each image each channel. Visual impression from RGB visualization seems that the relative strength between channels also depends on the experiment.",
      "votes": null
    },
    {
      "id": "621094",
      "postDate": "09/08/2019 07:41:32",
      "content": "<p>I guess any provided information could help to get a better result. However, I am also not sure how to use this <code>pixel_stats.csv</code>, or <code>positive/negative_treatment</code> examples. Still thinking about how I can \"merge\" this information into the training algorithm.​</p>",
      "rawMarkdown": "I guess any provided information could help to get a better result. However, I am also not sure how to use this `pixel_stats.csv`, or `positive/negative_treatment` examples. Still thinking about how I can \"merge\" this information into the training algorithm.​",
      "votes": null
    },
    {
      "id": "626154",
      "postDate": "09/13/2019 23:59:05",
      "content": "<p>I wonder, how are mean and std values from pixel_stats calculated? They look different from values one can get with numpy.</p>",
      "rawMarkdown": "I wonder, how are mean and std values from pixel_stats calculated? They look different from values one can get with numpy.",
      "votes": null
    },
    {
      "id": "626272",
      "postDate": "09/14/2019 05:27:11",
      "content": "<p>pixel_stats are computed in range [0, 255] while images are in [0, 1] (maybe depends on how you read?). I got almost the same mean and std after multiplying by 255.</p>\n\n<p>```\nid_code experiment  channel mean    std median  min max\n0   HEPG2-01_1_B02  1   71.063782   43.146240   67.0    7   255</p>\n\n<p>np.mean(img[0, ])*255  # for img.shape = [6, 512, 512]\n=&gt; 71.06378362\n```\n​​</p>",
      "rawMarkdown": "pixel_stats are computed in range [0, 255] while images are in [0, 1] (maybe depends on how you read?). I got almost the same mean and std after multiplying by 255.\n\n```\nid_code\texperiment\tchannel\tmean\tstd\tmedian\tmin\tmax\n0\tHEPG2-01_1_B02\t1\t71.063782\t43.146240\t67.0\t7\t255\n\nnp.mean(img[0, ])*255  # for img.shape = [6, 512, 512]\n=&gt; 71.06378362\n```\n​​",
      "votes": null
    },
    {
      "id": "626356",
      "postDate": "09/14/2019 07:54:40",
      "content": "<p>Thanks for reply, I've already got it, seems like I just messed with the indexes. But then, why even provide such stats, when they're so easy to calculate?</p>",
      "rawMarkdown": "Thanks for reply, I've already got it, seems like I just messed with the indexes. But then, why even provide such stats, when they're so easy to calculate?",
      "votes": null
    },
    {
      "id": "627417",
      "postDate": "09/15/2019 23:51:56",
      "content": "<p>I think I might be missing the core concept here. Using the pixel info seems like a good idea to help account for different bias, but wouldnt it be better to calculate the values when doing normalization after any augmentations</p>",
      "rawMarkdown": "I think I might be missing the core concept here. Using the pixel info seems like a good idea to help account for different bias, but wouldnt it be better to calculate the values when doing normalization after any augmentations",
      "votes": null
    },
    {
      "id": "627642",
      "postDate": "09/16/2019 07:53:33",
      "content": "<p>In fact, you can just calculate the same values yourself, in my experience with this case standardization matters</p>",
      "rawMarkdown": "In fact, you can just calculate the same values yourself, in my experience with this case standardization matters",
      "votes": null
    },
    {
      "id": "631262",
      "postDate": "09/21/2019 17:40:40",
      "content": "<p><a href=\"/cateek\">@cateek</a>  If possible can you tell or give hints on what standardization method is better here.</p>",
      "rawMarkdown": "cateek  If possible can you tell or give hints on what standardization method is better here.",
      "votes": null
    },
    {
      "id": "631287",
      "postDate": "09/21/2019 19:00:49",
      "content": "<p>I'm trying to figure that out myself actually) For now I'm just using tf.image.per_image_standardization on every channel</p>",
      "rawMarkdown": "I'm trying to figure that out myself actually) For now I'm just using tf.image.per_image_standardization on every channel",
      "votes": null
    },
    {
      "id": "631817",
      "postDate": "09/22/2019 18:22:51",
      "content": "<p>Thanks for letting me know. I've been using the per channel normalization and can't score more than 0.3 LB without leak with efficientnet with augmentation. Any pointers from your side.</p>",
      "rawMarkdown": "Thanks for letting me know. I've been using the per channel normalization and can't score more than 0.3 LB without leak with efficientnet with augmentation. Any pointers from your side.",
      "votes": null
    },
    {
      "id": "631890",
      "postDate": "09/22/2019 22:04:54",
      "content": "<p>It's really hard to give any advice without knowing any details.  I'm also using efficientnet, with 6 channel images and both sites. Augmentation improves score but not as standardization. </p>",
      "rawMarkdown": "It's really hard to give any advice without knowing any details.  I'm also using efficientnet, with 6 channel images and both sites. Augmentation improves score but not as standardization.",
      "votes": null
    },
    {
      "id": "632018",
      "postDate": "09/23/2019 05:58:42",
      "content": "<p>Thanks.I'll probably give one last shot at it.</p>",
      "rawMarkdown": "Thanks.I'll probably give one last shot at it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 619890,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "09/06/2019 17:57:44",
      "content": "<p>I think so. One of the themes of this competition is to predict robustly for data from different experiments (batch effect). I am not using pixel_stats.csv but computing a normalization factor for each image each channel. Visual impression from RGB visualization seems that the relative strength between channels also depends on the experiment.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 621094,
      "author_name": "purplejester",
      "author_url": "",
      "post_date": "09/08/2019 07:41:32",
      "content": "<p>I guess any provided information could help to get a better result. However, I am also not sure how to use this <code>pixel_stats.csv</code>, or <code>positive/negative_treatment</code> examples. Still thinking about how I can \"merge\" this information into the training algorithm.​</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 626154,
      "author_name": "cateek",
      "author_url": "",
      "post_date": "09/13/2019 23:59:05",
      "content": "<p>I wonder, how are mean and std values from pixel_stats calculated? They look different from values one can get with numpy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 626272,
          "author_name": "junkoda",
          "author_url": "",
          "post_date": "09/14/2019 05:27:11",
          "content": "<p>pixel_stats are computed in range [0, 255] while images are in [0, 1] (maybe depends on how you read?). I got almost the same mean and std after multiplying by 255.</p>\n\n<p>```\nid_code experiment  channel mean    std median  min max\n0   HEPG2-01_1_B02  1   71.063782   43.146240   67.0    7   255</p>\n\n<p>np.mean(img[0, ])*255  # for img.shape = [6, 512, 512]\n=&gt; 71.06378362\n```\n​​</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 626356,
          "author_name": "cateek",
          "author_url": "",
          "post_date": "09/14/2019 07:54:40",
          "content": "<p>Thanks for reply, I've already got it, seems like I just messed with the indexes. But then, why even provide such stats, when they're so easy to calculate?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 627417,
      "author_name": "governor",
      "author_url": "",
      "post_date": "09/15/2019 23:51:56",
      "content": "<p>I think I might be missing the core concept here. Using the pixel info seems like a good idea to help account for different bias, but wouldnt it be better to calculate the values when doing normalization after any augmentations</p>",
      "votes": null,
      "replies": [
        {
          "id": 627642,
          "author_name": "cateek",
          "author_url": "",
          "post_date": "09/16/2019 07:53:33",
          "content": "<p>In fact, you can just calculate the same values yourself, in my experience with this case standardization matters</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 631262,
          "author_name": "ekansh",
          "author_url": "",
          "post_date": "09/21/2019 17:40:40",
          "content": "<p><a href=\"/cateek\">@cateek</a>  If possible can you tell or give hints on what standardization method is better here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 631287,
          "author_name": "cateek",
          "author_url": "",
          "post_date": "09/21/2019 19:00:49",
          "content": "<p>I'm trying to figure that out myself actually) For now I'm just using tf.image.per_image_standardization on every channel</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 631817,
          "author_name": "ekansh",
          "author_url": "",
          "post_date": "09/22/2019 18:22:51",
          "content": "<p>Thanks for letting me know. I've been using the per channel normalization and can't score more than 0.3 LB without leak with efficientnet with augmentation. Any pointers from your side.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 631890,
          "author_name": "cateek",
          "author_url": "",
          "post_date": "09/22/2019 22:04:54",
          "content": "<p>It's really hard to give any advice without knowing any details.  I'm also using efficientnet, with 6 channel images and both sites. Augmentation improves score but not as standardization. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 632018,
          "author_name": "ekansh",
          "author_url": "",
          "post_date": "09/23/2019 05:58:42",
          "content": "<p>Thanks.I'll probably give one last shot at it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "619623": "Do we need to normalize EACH of our image with the specific mean and std from the pixel_stats.csv?\n\nUsually I will normalize the whole dataset by pass the same mean and std by torchvision.transforms in pytorch. I am newbie and first time for me to come into this situation.\n\nThanks",
    "619890": "I think so. One of the themes of this competition is to predict robustly for data from different experiments (batch effect). I am not using pixel_stats.csv but computing a normalization factor for each image each channel. Visual impression from RGB visualization seems that the relative strength between channels also depends on the experiment.",
    "621094": "I guess any provided information could help to get a better result. However, I am also not sure how to use this `pixel_stats.csv`, or `positive/negative_treatment` examples. Still thinking about how I can \"merge\" this information into the training algorithm.​",
    "626154": "I wonder, how are mean and std values from pixel_stats calculated? They look different from values one can get with numpy.",
    "626272": "pixel_stats are computed in range [0, 255] while images are in [0, 1] (maybe depends on how you read?). I got almost the same mean and std after multiplying by 255.\n\n```\nid_code\texperiment\tchannel\tmean\tstd\tmedian\tmin\tmax\n0\tHEPG2-01_1_B02\t1\t71.063782\t43.146240\t67.0\t7\t255\n\nnp.mean(img[0, ])*255  # for img.shape = [6, 512, 512]\n=&gt; 71.06378362\n```\n​​",
    "626356": "Thanks for reply, I've already got it, seems like I just messed with the indexes. But then, why even provide such stats, when they're so easy to calculate?",
    "627417": "I think I might be missing the core concept here. Using the pixel info seems like a good idea to help account for different bias, but wouldnt it be better to calculate the values when doing normalization after any augmentations",
    "627642": "In fact, you can just calculate the same values yourself, in my experience with this case standardization matters",
    "631262": "cateek  If possible can you tell or give hints on what standardization method is better here.",
    "631287": "I'm trying to figure that out myself actually) For now I'm just using tf.image.per_image_standardization on every channel",
    "631817": "Thanks for letting me know. I've been using the per channel normalization and can't score more than 0.3 LB without leak with efficientnet with augmentation. Any pointers from your side.",
    "631890": "It's really hard to give any advice without knowing any details.  I'm also using efficientnet, with 6 channel images and both sites. Augmentation improves score but not as standardization.",
    "632018": "Thanks.I'll probably give one last shot at it."
  },
  "source": "meta"
}