{
  "id": 128684,
  "title": "How to normalize the data when using pre-trained model",
  "url": "/competitions/bengaliai-cv19/discussion/128684",
  "author_name": "",
  "post_date": "2020-02-02T14:11:46.455131600Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi, all.</p>\n\n<p>I'm curious that when using pre-trained models, do you normalize the data with general value (mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]  for imageNet), or do you calculate the whole training set to get its own mean and std? (it's only one channel BTW).</p>\n\n<p>How's the performance differs between them?</p>",
  "messages": [
    {
      "id": "735069",
      "postDate": "02/02/2020 14:11:46",
      "content": "<p>Hi, all.</p>\n\n<p>I'm curious that when using pre-trained models, do you normalize the data with general value (mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]  for imageNet), or do you calculate the whole training set to get its own mean and std? (it's only one channel BTW).</p>\n\n<p>How's the performance differs between them?</p>",
      "rawMarkdown": "Hi, all.\n\nI'm curious that when using pre-trained models, do you normalize the data with general value (mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]  for imageNet), or do you calculate the whole training set to get its own mean and std? (it's only one channel BTW).\n\nHow's the performance differs between them?",
      "votes": null
    },
    {
      "id": "735107",
      "postDate": "02/02/2020 15:51:29",
      "content": "<p>As we don't have access to the test data, your best bet would be to normalize with mean and std from training set. </p>\n\n<p>IMO anything close to standard distribution works fine. \nEven a (0, 1) distribution works fine with most NN classifiers. They only dislike large outliers.</p>",
      "rawMarkdown": "As we don't have access to the test data, your best bet would be to normalize with mean and std from training set. \n\nIMO anything close to standard distribution works fine. \nEven a (0, 1) distribution works fine with most NN classifiers. They only dislike large outliers.",
      "votes": null
    },
    {
      "id": "735139",
      "postDate": "02/02/2020 16:34:10",
      "content": "<p>I use single channel image and only divide by 255. Tbh, how you normalize the data will have very little effect on the model performance</p>",
      "rawMarkdown": "I use single channel image and only divide by 255. Tbh, how you normalize the data will have very little effect on the model performance",
      "votes": null
    },
    {
      "id": "735396",
      "postDate": "02/03/2020 01:14:49",
      "content": "<p>Thanks for replying! as you mentioned, in my experiment it did not change that much, I'll focus on other issues to improve my model.</p>",
      "rawMarkdown": "Thanks for replying! as you mentioned, in my experiment it did not change that much, I'll focus on other issues to improve my model.",
      "votes": null
    },
    {
      "id": "735398",
      "postDate": "02/03/2020 01:17:22",
      "content": "<p>Thanks for your reply!</p>",
      "rawMarkdown": "Thanks for your reply!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 735107,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "02/02/2020 15:51:29",
      "content": "<p>As we don't have access to the test data, your best bet would be to normalize with mean and std from training set. </p>\n\n<p>IMO anything close to standard distribution works fine. \nEven a (0, 1) distribution works fine with most NN classifiers. They only dislike large outliers.</p>",
      "votes": null,
      "replies": [
        {
          "id": 735398,
          "author_name": "ccchang801023",
          "author_url": "",
          "post_date": "02/03/2020 01:17:22",
          "content": "<p>Thanks for your reply!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 735139,
      "author_name": "andy2709",
      "author_url": "",
      "post_date": "02/02/2020 16:34:10",
      "content": "<p>I use single channel image and only divide by 255. Tbh, how you normalize the data will have very little effect on the model performance</p>",
      "votes": null,
      "replies": [
        {
          "id": 735396,
          "author_name": "ccchang801023",
          "author_url": "",
          "post_date": "02/03/2020 01:14:49",
          "content": "<p>Thanks for replying! as you mentioned, in my experiment it did not change that much, I'll focus on other issues to improve my model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "735069": "Hi, all.\n\nI'm curious that when using pre-trained models, do you normalize the data with general value (mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]  for imageNet), or do you calculate the whole training set to get its own mean and std? (it's only one channel BTW).\n\nHow's the performance differs between them?",
    "735107": "As we don't have access to the test data, your best bet would be to normalize with mean and std from training set. \n\nIMO anything close to standard distribution works fine. \nEven a (0, 1) distribution works fine with most NN classifiers. They only dislike large outliers.",
    "735139": "I use single channel image and only divide by 255. Tbh, how you normalize the data will have very little effect on the model performance",
    "735396": "Thanks for replying! as you mentioned, in my experiment it did not change that much, I'll focus on other issues to improve my model.",
    "735398": "Thanks for your reply!"
  },
  "source": "meta"
}