{
  "id": 138653,
  "title": "Should we be normalising inputs?",
  "url": "/competitions/flower-classification-with-tpus/discussion/138653",
  "author_name": "",
  "post_date": "2020-03-25T20:37:28.108359400Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I noticed that the prominent public notebooks here don't normalise the inputs (that is, they don't subtract channel-wise mean from the date before fitting or predicting). I gave it a try to see if it would make a difference (with an EfficientNet-B7 base) and it actually didn't change anything! I tried two variations of my experiment: one in which I freeze all EffNet weights but just train top, and the other in which I train everything together. There was no significant difference in the outcome.</p>\n\n<p>That's a bit confusing to me because I would think that one should follow the same data preparation is in the pretraining. So if the pretraining used channel-wise mean subtraction on the inputs, the weights of the conv layers will turn out to be different than if it didn't subtract channel-wise mean.</p>\n\n<p>Is something going on under the hood that I'm missing? Or is my last paragraph just wrong?</p>",
  "messages": [
    {
      "id": "786356",
      "postDate": "03/25/2020 20:37:28",
      "content": "<p>I noticed that the prominent public notebooks here don't normalise the inputs (that is, they don't subtract channel-wise mean from the date before fitting or predicting). I gave it a try to see if it would make a difference (with an EfficientNet-B7 base) and it actually didn't change anything! I tried two variations of my experiment: one in which I freeze all EffNet weights but just train top, and the other in which I train everything together. There was no significant difference in the outcome.</p>\n\n<p>That's a bit confusing to me because I would think that one should follow the same data preparation is in the pretraining. So if the pretraining used channel-wise mean subtraction on the inputs, the weights of the conv layers will turn out to be different than if it didn't subtract channel-wise mean.</p>\n\n<p>Is something going on under the hood that I'm missing? Or is my last paragraph just wrong?</p>",
      "rawMarkdown": "I noticed that the prominent public notebooks here don't normalise the inputs (that is, they don't subtract channel-wise mean from the date before fitting or predicting). I gave it a try to see if it would make a difference (with an EfficientNet-B7 base) and it actually didn't change anything! I tried two variations of my experiment: one in which I freeze all EffNet weights but just train top, and the other in which I train everything together. There was no significant difference in the outcome.\n\nThat's a bit confusing to me because I would think that one should follow the same data preparation is in the pretraining. So if the pretraining used channel-wise mean subtraction on the inputs, the weights of the conv layers will turn out to be different than if it didn't subtract channel-wise mean.\n\nIs something going on under the hood that I'm missing? Or is my last paragraph just wrong?",
      "votes": null
    },
    {
      "id": "786475",
      "postDate": "03/25/2020 23:25:50",
      "content": "<p>Hello Alexander\nI am also interested in this issue and want to hear explanation from experienced people.\nI can only answer your second question about freezing backbone layers. In my practice unfreezing all layers gives me always better result than freezing backbone. Also many people uses strategies like freezing backbone layers train for some epochs than unfreeze and train again (finetuning). If you tried unfreezing and it does not give you better result that may be because optimal learning rate in \"finetuning stage\"(training unfreezed backbone) are usually  ten times less than freezed.</p>",
      "rawMarkdown": "Hello Alexander\nI am also interested in this issue and want to hear explanation from experienced people.\nI can only answer your second question about freezing backbone layers. In my practice unfreezing all layers gives me always better result than freezing backbone. Also many people uses strategies like freezing backbone layers train for some epochs than unfreeze and train again (finetuning). If you tried unfreezing and it does not give you better result that may be because optimal learning rate in \"finetuning stage\"(training unfreezed backbone) are usually  ten times less than freezed.",
      "votes": null
    },
    {
      "id": "786519",
      "postDate": "03/26/2020 00:22:47",
      "content": "<p>Two possible reasons:\n- image data are fairly normalized already: all values in [0.0, 1.0], mean not far from 0.5. The 0.5 offset can easily be fixed by the bias of the first layer.\n- the use of batch normalization in most modern neural networks</p>",
      "rawMarkdown": "Two possible reasons:\n- image data are fairly normalized already: all values in [0.0, 1.0], mean not far from 0.5. The 0.5 offset can easily be fixed by the bias of the first layer.\n- the use of batch normalization in most modern neural networks",
      "votes": null
    },
    {
      "id": "786762",
      "postDate": "03/26/2020 07:21:49",
      "content": "<p>Hi. Thanks for that. I did the experiment of freeze vs not freeze because I thought the former would give a more obvious difference in the normalise inputs vs don't normalise inputs dimension. Please upvote the topic of you find it useful - that way it will get more visibility and is likely to get an answer.</p>",
      "rawMarkdown": "Hi. Thanks for that. I did the experiment of freeze vs not freeze because I thought the former would give a more obvious difference in the normalise inputs vs don't normalise inputs dimension. Please upvote the topic of you find it useful - that way it will get more visibility and is likely to get an answer.",
      "votes": null
    },
    {
      "id": "786902",
      "postDate": "03/26/2020 10:18:44",
      "content": "<p>Thanks! Glad you answered, because now I know that the tf/tpu implementation is not doing stuff under the hood. And yes I suppose the batch-norm might come before the first relu which would make it invariant to a constant shift (will verify). Although if the relu came before the batch norm maybe it would be an issue.</p>",
      "rawMarkdown": "Thanks! Glad you answered, because now I know that the tf/tpu implementation is not doing stuff under the hood. And yes I suppose the batch-norm might come before the first relu which would make it invariant to a constant shift (will verify). Although if the relu came before the batch norm maybe it would be an issue.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 786475,
      "author_name": "xarshila",
      "author_url": "",
      "post_date": "03/25/2020 23:25:50",
      "content": "<p>Hello Alexander\nI am also interested in this issue and want to hear explanation from experienced people.\nI can only answer your second question about freezing backbone layers. In my practice unfreezing all layers gives me always better result than freezing backbone. Also many people uses strategies like freezing backbone layers train for some epochs than unfreeze and train again (finetuning). If you tried unfreezing and it does not give you better result that may be because optimal learning rate in \"finetuning stage\"(training unfreezed backbone) are usually  ten times less than freezed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 786762,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "03/26/2020 07:21:49",
          "content": "<p>Hi. Thanks for that. I did the experiment of freeze vs not freeze because I thought the former would give a more obvious difference in the normalise inputs vs don't normalise inputs dimension. Please upvote the topic of you find it useful - that way it will get more visibility and is likely to get an answer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 786519,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/26/2020 00:22:47",
      "content": "<p>Two possible reasons:\n- image data are fairly normalized already: all values in [0.0, 1.0], mean not far from 0.5. The 0.5 offset can easily be fixed by the bias of the first layer.\n- the use of batch normalization in most modern neural networks</p>",
      "votes": null,
      "replies": [
        {
          "id": 786902,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "03/26/2020 10:18:44",
          "content": "<p>Thanks! Glad you answered, because now I know that the tf/tpu implementation is not doing stuff under the hood. And yes I suppose the batch-norm might come before the first relu which would make it invariant to a constant shift (will verify). Although if the relu came before the batch norm maybe it would be an issue.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "786356": "I noticed that the prominent public notebooks here don't normalise the inputs (that is, they don't subtract channel-wise mean from the date before fitting or predicting). I gave it a try to see if it would make a difference (with an EfficientNet-B7 base) and it actually didn't change anything! I tried two variations of my experiment: one in which I freeze all EffNet weights but just train top, and the other in which I train everything together. There was no significant difference in the outcome.\n\nThat's a bit confusing to me because I would think that one should follow the same data preparation is in the pretraining. So if the pretraining used channel-wise mean subtraction on the inputs, the weights of the conv layers will turn out to be different than if it didn't subtract channel-wise mean.\n\nIs something going on under the hood that I'm missing? Or is my last paragraph just wrong?",
    "786475": "Hello Alexander\nI am also interested in this issue and want to hear explanation from experienced people.\nI can only answer your second question about freezing backbone layers. In my practice unfreezing all layers gives me always better result than freezing backbone. Also many people uses strategies like freezing backbone layers train for some epochs than unfreeze and train again (finetuning). If you tried unfreezing and it does not give you better result that may be because optimal learning rate in \"finetuning stage\"(training unfreezed backbone) are usually  ten times less than freezed.",
    "786519": "Two possible reasons:\n- image data are fairly normalized already: all values in [0.0, 1.0], mean not far from 0.5. The 0.5 offset can easily be fixed by the bias of the first layer.\n- the use of batch normalization in most modern neural networks",
    "786762": "Hi. Thanks for that. I did the experiment of freeze vs not freeze because I thought the former would give a more obvious difference in the normalise inputs vs don't normalise inputs dimension. Please upvote the topic of you find it useful - that way it will get more visibility and is likely to get an answer.",
    "786902": "Thanks! Glad you answered, because now I know that the tf/tpu implementation is not doing stuff under the hood. And yes I suppose the batch-norm might come before the first relu which would make it invariant to a constant shift (will verify). Although if the relu came before the batch norm maybe it would be an issue."
  },
  "source": "meta"
}