{
  "id": 101619,
  "title": "Image data normalization",
  "url": "/competitions/aptos2019-blindness-detection/discussion/101619",
  "author_name": "",
  "post_date": "2019-07-27T08:07:32.218421400Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I use image data normalization for training my model:</p>\n\n<p><code>train_datagen =  ImageDataGenerator(\n            featurewise_center=True,\n            featurewise_std_normalization=True\n            ....\ntrain_datagen.fit(x_train)\n</code></p>\n\n<p>I wonder what is the best practice, when I use several datasets to train my model.\nFirst I train with the the training <strong>data from the previous competition</strong>\nThen I continue training with the <strong>new dataset</strong>\nAt last I make predictions of for the <strong>test set</strong>.</p>\n\n<p><strong>The mean and std values for these datasets are very different.</strong>\nShould I use the same normalization for each step?\nOr should I always use the mean and std of the current dataset?</p>\n\n<p>What is the theory behind this?\nCan anyone help?</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "585294",
      "postDate": "07/27/2019 08:07:32",
      "content": "<p>Hi,</p>\n\n<p>I use image data normalization for training my model:</p>\n\n<p><code>train_datagen =  ImageDataGenerator(\n            featurewise_center=True,\n            featurewise_std_normalization=True\n            ....\ntrain_datagen.fit(x_train)\n</code></p>\n\n<p>I wonder what is the best practice, when I use several datasets to train my model.\nFirst I train with the the training <strong>data from the previous competition</strong>\nThen I continue training with the <strong>new dataset</strong>\nAt last I make predictions of for the <strong>test set</strong>.</p>\n\n<p><strong>The mean and std values for these datasets are very different.</strong>\nShould I use the same normalization for each step?\nOr should I always use the mean and std of the current dataset?</p>\n\n<p>What is the theory behind this?\nCan anyone help?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi,\n\nI use image data normalization for training my model:\n\n`train_datagen =  ImageDataGenerator(\n            featurewise_center=True,\n            featurewise_std_normalization=True\n            ....\ntrain_datagen.fit(x_train)\n`\n\nI wonder what is the best practice, when I use several datasets to train my model.\nFirst I train with the the training **data from the previous competition**\nThen I continue training with the **new dataset**\nAt last I make predictions of for the **test set**.\n\n**The mean and std values for these datasets are very different.**\nShould I use the same normalization for each step?\nOr should I always use the mean and std of the current dataset?\n\nWhat is the theory behind this?\nCan anyone help?\n\nThanks",
      "votes": null
    },
    {
      "id": "588916",
      "postDate": "07/31/2019 08:13:18",
      "content": "<p>I have tried samplewise_center/samplewise_std_normalization when training, but it makes model worse, the kappa score is always 0 approximately in the whole training steps,  it seems that the model(i use densenet) expect an input of certain form.\nBut I have not tried featurewise_center/featurewise_std_normalization, dose it work normally?</p>",
      "rawMarkdown": "I have tried samplewise_center/samplewise_std_normalization when training, but it makes model worse, the kappa score is always 0 approximately in the whole training steps,  it seems that the model(i use densenet) expect an input of certain form.\nBut I have not tried featurewise_center/featurewise_std_normalization, dose it work normally?",
      "votes": null
    },
    {
      "id": "590559",
      "postDate": "08/02/2019 10:38:34",
      "content": "<p>Hi, I have taken help from the following link for image data normalization:\n<a href=\"https://machinelearningmastery.com/how-to-manually-scale-image-pixel-data-for-deep-learning/\">https://machinelearningmastery.com/how-to-manually-scale-image-pixel-data-for-deep-learning/</a>,\nyou can check it</p>",
      "rawMarkdown": "Hi, I have taken help from the following link for image data normalization:\nhttps://machinelearningmastery.com/how-to-manually-scale-image-pixel-data-for-deep-learning/,\nyou can check it",
      "votes": null
    },
    {
      "id": "590583",
      "postDate": "08/02/2019 11:11:53",
      "content": "<p>It does work, but I'm not sure how to normalize 3 different data sets...\nNow I just scale all with 1/255</p>",
      "rawMarkdown": "It does work, but I'm not sure how to normalize 3 different data sets...\nNow I just scale all with 1/255",
      "votes": null
    },
    {
      "id": "591797",
      "postDate": "08/04/2019 09:14:05",
      "content": "<p>IMO, the purpose of normalization is : let the model see the same value range of the pictures no matter training , inferring or using old dataset. So, i think you should using each dataset's own mean/std to do normalization.\nBTW, i have a question: after normalization, the pixel value type will be float, including negative values, but in many image lib, pixel value range is [0, 1] if value type is float, if you just scale all with 1/255, how about negative values, which means \"black\" in those image lib.</p>",
      "rawMarkdown": "IMO, the purpose of normalization is : let the model see the same value range of the pictures no matter training , inferring or using old dataset. So, i think you should using each dataset's own mean/std to do normalization.\nBTW, i have a question: after normalization, the pixel value type will be float, including negative values, but in many image lib, pixel value range is [0, 1] if value type is float, if you just scale all with 1/255, how about negative values, which means \"black\" in those image lib.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 588916,
      "author_name": "frank518",
      "author_url": "",
      "post_date": "07/31/2019 08:13:18",
      "content": "<p>I have tried samplewise_center/samplewise_std_normalization when training, but it makes model worse, the kappa score is always 0 approximately in the whole training steps,  it seems that the model(i use densenet) expect an input of certain form.\nBut I have not tried featurewise_center/featurewise_std_normalization, dose it work normally?</p>",
      "votes": null,
      "replies": [
        {
          "id": 590583,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "08/02/2019 11:11:53",
          "content": "<p>It does work, but I'm not sure how to normalize 3 different data sets...\nNow I just scale all with 1/255</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 591797,
          "author_name": "frank518",
          "author_url": "",
          "post_date": "08/04/2019 09:14:05",
          "content": "<p>IMO, the purpose of normalization is : let the model see the same value range of the pictures no matter training , inferring or using old dataset. So, i think you should using each dataset's own mean/std to do normalization.\nBTW, i have a question: after normalization, the pixel value type will be float, including negative values, but in many image lib, pixel value range is [0, 1] if value type is float, if you just scale all with 1/255, how about negative values, which means \"black\" in those image lib.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 590559,
      "author_name": "nanditab35",
      "author_url": "",
      "post_date": "08/02/2019 10:38:34",
      "content": "<p>Hi, I have taken help from the following link for image data normalization:\n<a href=\"https://machinelearningmastery.com/how-to-manually-scale-image-pixel-data-for-deep-learning/\">https://machinelearningmastery.com/how-to-manually-scale-image-pixel-data-for-deep-learning/</a>,\nyou can check it</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "585294": "Hi,\n\nI use image data normalization for training my model:\n\n`train_datagen =  ImageDataGenerator(\n            featurewise_center=True,\n            featurewise_std_normalization=True\n            ....\ntrain_datagen.fit(x_train)\n`\n\nI wonder what is the best practice, when I use several datasets to train my model.\nFirst I train with the the training **data from the previous competition**\nThen I continue training with the **new dataset**\nAt last I make predictions of for the **test set**.\n\n**The mean and std values for these datasets are very different.**\nShould I use the same normalization for each step?\nOr should I always use the mean and std of the current dataset?\n\nWhat is the theory behind this?\nCan anyone help?\n\nThanks",
    "588916": "I have tried samplewise_center/samplewise_std_normalization when training, but it makes model worse, the kappa score is always 0 approximately in the whole training steps,  it seems that the model(i use densenet) expect an input of certain form.\nBut I have not tried featurewise_center/featurewise_std_normalization, dose it work normally?",
    "590559": "Hi, I have taken help from the following link for image data normalization:\nhttps://machinelearningmastery.com/how-to-manually-scale-image-pixel-data-for-deep-learning/,\nyou can check it",
    "590583": "It does work, but I'm not sure how to normalize 3 different data sets...\nNow I just scale all with 1/255",
    "591797": "IMO, the purpose of normalization is : let the model see the same value range of the pictures no matter training , inferring or using old dataset. So, i think you should using each dataset's own mean/std to do normalization.\nBTW, i have a question: after normalization, the pixel value type will be float, including negative values, but in many image lib, pixel value range is [0, 1] if value type is float, if you just scale all with 1/255, how about negative values, which means \"black\" in those image lib."
  },
  "source": "meta"
}