{
  "id": 68934,
  "title": "Image preprocessing scheme when using transfer learning from ImageNet model??",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/68934",
  "author_name": "",
  "post_date": "2018-10-18T15:19:48.450748900Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I saw that majority of the kernels DID NOT contain image preprocessing elements in them (i.e. zero-center the images, rescale images to have (0,1) mean and std etc.). Instead, most of them just divide the RCB images by 255 to make them range from 0 to 1. I was wondering how will different image preprocessing schemes impact your final prediction results/training performance, given that the training is started from a pre-trained model based on pre-processed ImageNet images.</p>\n\n<p>If one is using the pre-trained model from Keras applications, they have their built-in preprocessing function with the following modes (options):</p>\n\n<pre><code>    mode: One of \"caffe\", \"tf\" or \"torch\".\n        - caffe: will convert the images from RGB to BGR,\n            then will zero-center each color channel with\n            respect to the ImageNet dataset,\n            without scaling.\n        - tf: will scale pixels between -1 and 1,\n            sample-wise.\n        - torch: will scale pixels between 0 and 1 and then\n            will normalize each channel with respect to the\n            ImageNet dataset.\n</code></pre>\n\n<p>If anyone has experiment on this issue before, will you please share your insight? Otherwise, I will do some little experiment on myself...</p>\n\n<p>Thanks in advance!!</p>",
  "messages": [
    {
      "id": "406056",
      "postDate": "10/18/2018 15:19:48",
      "content": "<p>Hi,</p>\n\n<p>I saw that majority of the kernels DID NOT contain image preprocessing elements in them (i.e. zero-center the images, rescale images to have (0,1) mean and std etc.). Instead, most of them just divide the RCB images by 255 to make them range from 0 to 1. I was wondering how will different image preprocessing schemes impact your final prediction results/training performance, given that the training is started from a pre-trained model based on pre-processed ImageNet images.</p>\n\n<p>If one is using the pre-trained model from Keras applications, they have their built-in preprocessing function with the following modes (options):</p>\n\n<pre><code>    mode: One of \"caffe\", \"tf\" or \"torch\".\n        - caffe: will convert the images from RGB to BGR,\n            then will zero-center each color channel with\n            respect to the ImageNet dataset,\n            without scaling.\n        - tf: will scale pixels between -1 and 1,\n            sample-wise.\n        - torch: will scale pixels between 0 and 1 and then\n            will normalize each channel with respect to the\n            ImageNet dataset.\n</code></pre>\n\n<p>If anyone has experiment on this issue before, will you please share your insight? Otherwise, I will do some little experiment on myself...</p>\n\n<p>Thanks in advance!!</p>",
      "rawMarkdown": "Hi,\n\nI saw that majority of the kernels DID NOT contain image preprocessing elements in them (i.e. zero-center the images, rescale images to have (0,1) mean and std etc.). Instead, most of them just divide the RCB images by 255 to make them range from 0 to 1. I was wondering how will different image preprocessing schemes impact your final prediction results/training performance, given that the training is started from a pre-trained model based on pre-processed ImageNet images.\n\nIf one is using the pre-trained model from Keras applications, they have their built-in preprocessing function with the following modes (options):\n\n        mode: One of \"caffe\", \"tf\" or \"torch\".\n            - caffe: will convert the images from RGB to BGR,\n                then will zero-center each color channel with\n                respect to the ImageNet dataset,\n                without scaling.\n            - tf: will scale pixels between -1 and 1,\n                sample-wise.\n            - torch: will scale pixels between 0 and 1 and then\n                will normalize each channel with respect to the\n                ImageNet dataset.\n\nIf anyone has experiment on this issue before, will you please share your insight? Otherwise, I will do some little experiment on myself...\n\nThanks in advance!!",
      "votes": null
    },
    {
      "id": "406348",
      "postDate": "10/19/2018 05:22:00",
      "content": "<p>I use the keras.imagenet_utils.preprocess_input. Previously I would only divide by 255. It trains a little faster with preprocess_input, but not significantly. If I had to put a number on it I would say 5-10% faster training. I've not measured it specifically. Using Keras it was easy to stick into my data generator:</p>\n\n<pre>train_args = dict(featurewise_center = False, \n                samplewise_center = False,\n                horizontal_flip = True, \n                vertical_flip = True,\n                fill_mode = 'reflect',\n                data_format = 'channels_last',\n                preprocessing_function = preprocess_input)\n</pre>",
      "rawMarkdown": "I use the keras.imagenet_utils.preprocess_input. Previously I would only divide by 255. It trains a little faster with preprocess_input, but not significantly. If I had to put a number on it I would say 5-10% faster training. I've not measured it specifically. Using Keras it was easy to stick into my data generator:\n\n<pre>train_args = dict(featurewise_center = False, \n                samplewise_center = False,\n                horizontal_flip = True, \n                vertical_flip = True,\n                fill_mode = 'reflect',\n                data_format = 'channels_last',\n                preprocessing_function = preprocess_input)\n</pre>",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 406348,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "10/19/2018 05:22:00",
      "content": "<p>I use the keras.imagenet_utils.preprocess_input. Previously I would only divide by 255. It trains a little faster with preprocess_input, but not significantly. If I had to put a number on it I would say 5-10% faster training. I've not measured it specifically. Using Keras it was easy to stick into my data generator:</p>\n\n<pre>train_args = dict(featurewise_center = False, \n                samplewise_center = False,\n                horizontal_flip = True, \n                vertical_flip = True,\n                fill_mode = 'reflect',\n                data_format = 'channels_last',\n                preprocessing_function = preprocess_input)\n</pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "406056": "Hi,\n\nI saw that majority of the kernels DID NOT contain image preprocessing elements in them (i.e. zero-center the images, rescale images to have (0,1) mean and std etc.). Instead, most of them just divide the RCB images by 255 to make them range from 0 to 1. I was wondering how will different image preprocessing schemes impact your final prediction results/training performance, given that the training is started from a pre-trained model based on pre-processed ImageNet images.\n\nIf one is using the pre-trained model from Keras applications, they have their built-in preprocessing function with the following modes (options):\n\n        mode: One of \"caffe\", \"tf\" or \"torch\".\n            - caffe: will convert the images from RGB to BGR,\n                then will zero-center each color channel with\n                respect to the ImageNet dataset,\n                without scaling.\n            - tf: will scale pixels between -1 and 1,\n                sample-wise.\n            - torch: will scale pixels between 0 and 1 and then\n                will normalize each channel with respect to the\n                ImageNet dataset.\n\nIf anyone has experiment on this issue before, will you please share your insight? Otherwise, I will do some little experiment on myself...\n\nThanks in advance!!",
    "406348": "I use the keras.imagenet_utils.preprocess_input. Previously I would only divide by 255. It trains a little faster with preprocess_input, but not significantly. If I had to put a number on it I would say 5-10% faster training. I've not measured it specifically. Using Keras it was easy to stick into my data generator:\n\n<pre>train_args = dict(featurewise_center = False, \n                samplewise_center = False,\n                horizontal_flip = True, \n                vertical_flip = True,\n                fill_mode = 'reflect',\n                data_format = 'channels_last',\n                preprocessing_function = preprocess_input)\n</pre>"
  },
  "source": "meta"
}