{
  "id": 120379,
  "title": "Filter Response Normalization (FRN)",
  "url": "/competitions/pku-autonomous-driving/discussion/120379",
  "author_name": "",
  "post_date": "2019-12-05T16:01:03.276039900Z",
  "votes": 17,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Last week Google AI released research on a new normalization layer called<a href=\"https://arxiv.org/pdf/1911.09737.pdf\"> \"Filter Response Normalization\" (FRN) together with the \"Thresholded Linear Unit\" (TLU) activation</a>. They observed considerable gains in accuracy across all batch size, but I wanted to see for myself if it would work well on other data science problems. If the claims are true then it would make Batch Normalization, Group Normalization and Layer Normalization kind of obsolete.</p>\n\n<p>FRN</p>\n\n<p><img src=\"https://deeplearn.org/arxiv_files/1911.09737v1/x1.png\" alt=\"\"></p>\n\n<p>Image: A comparison of FRN with BN, BRN and GN. <a href=\"https://deeplearn.org/arxiv/104635/filter-response-normalization-layer%3a-eliminating-batch-dependence-in-the-training-of-deep-neural-networks\">Source</a></p>\n\n<p>I'm still experimenting with it, but just uploaded <a href=\"https://github.com/CarloLepelaars/filter_response_normalization_keras\">an implementation of this on Github</a>. I tested it on EfficientNetB5 with a small batch size (4) and it seems to work well. However, I still have to do some more experiments to see if it is beneficial compared to group normalization or layer normalization with small batch sizes.</p>\n\n<p>Please let me know here what your experiences are and if you see gains in performance from replacing your batch normalization and activation layers with FRN + TLU. Curious to see how well this research works in the real world!</p>\n\n<p>Paper:\n<a href=\"https://arxiv.org/pdf/1911.09737.pdf\">https://arxiv.org/pdf/1911.09737.pdf</a></p>\n\n<p>Repository with implementation using Tensorflow/Keras:\n<a href=\"https://github.com/CarloLepelaars/filter_response_normalization_keras\">https://github.com/CarloLepelaars/filter_response_normalization_keras</a></p>",
  "messages": [
    {
      "id": "688459",
      "postDate": "12/05/2019 16:01:03",
      "content": "<p>Last week Google AI released research on a new normalization layer called<a href=\"https://arxiv.org/pdf/1911.09737.pdf\"> \"Filter Response Normalization\" (FRN) together with the \"Thresholded Linear Unit\" (TLU) activation</a>. They observed considerable gains in accuracy across all batch size, but I wanted to see for myself if it would work well on other data science problems. If the claims are true then it would make Batch Normalization, Group Normalization and Layer Normalization kind of obsolete.</p>\n\n<p>FRN</p>\n\n<p><img src=\"https://deeplearn.org/arxiv_files/1911.09737v1/x1.png\" alt=\"\"></p>\n\n<p>Image: A comparison of FRN with BN, BRN and GN. <a href=\"https://deeplearn.org/arxiv/104635/filter-response-normalization-layer%3a-eliminating-batch-dependence-in-the-training-of-deep-neural-networks\">Source</a></p>\n\n<p>I'm still experimenting with it, but just uploaded <a href=\"https://github.com/CarloLepelaars/filter_response_normalization_keras\">an implementation of this on Github</a>. I tested it on EfficientNetB5 with a small batch size (4) and it seems to work well. However, I still have to do some more experiments to see if it is beneficial compared to group normalization or layer normalization with small batch sizes.</p>\n\n<p>Please let me know here what your experiences are and if you see gains in performance from replacing your batch normalization and activation layers with FRN + TLU. Curious to see how well this research works in the real world!</p>\n\n<p>Paper:\n<a href=\"https://arxiv.org/pdf/1911.09737.pdf\">https://arxiv.org/pdf/1911.09737.pdf</a></p>\n\n<p>Repository with implementation using Tensorflow/Keras:\n<a href=\"https://github.com/CarloLepelaars/filter_response_normalization_keras\">https://github.com/CarloLepelaars/filter_response_normalization_keras</a></p>",
      "rawMarkdown": "Last week Google AI released research on a new normalization layer called[ \"Filter Response Normalization\" (FRN) together with the \"Thresholded Linear Unit\" (TLU) activation](https://arxiv.org/pdf/1911.09737.pdf). They observed considerable gains in accuracy across all batch size, but I wanted to see for myself if it would work well on other data science problems. If the claims are true then it would make Batch Normalization, Group Normalization and Layer Normalization kind of obsolete.\n\nFRN\n\n![](https://deeplearn.org/arxiv_files/1911.09737v1/x1.png)\n\nImage: A comparison of FRN with BN, BRN and GN. [Source](https://deeplearn.org/arxiv/104635/filter-response-normalization-layer:-eliminating-batch-dependence-in-the-training-of-deep-neural-networks)\n\nI'm still experimenting with it, but just uploaded [an implementation of this on Github](https://github.com/CarloLepelaars/filter_response_normalization_keras). I tested it on EfficientNetB5 with a small batch size (4) and it seems to work well. However, I still have to do some more experiments to see if it is beneficial compared to group normalization or layer normalization with small batch sizes.\n\nPlease let me know here what your experiences are and if you see gains in performance from replacing your batch normalization and activation layers with FRN + TLU. Curious to see how well this research works in the real world!\n\nPaper:\nhttps://arxiv.org/pdf/1911.09737.pdf\n\nRepository with implementation using Tensorflow/Keras:\nhttps://github.com/CarloLepelaars/filter_response_normalization_keras",
      "votes": null
    },
    {
      "id": "688461",
      "postDate": "12/05/2019 16:02:34",
      "content": "<p>Quick code snippet:</p>\n\n<p>```\nfrom keras import backend as K\nfrom keras.engine import Layer\nfrom keras import initializers, regularizers, constraints</p>\n\n<p>class FilterResponseNormalization(Layer):\n    \"\"\"\n    Implementation of the Filter Response Normalization (FRN) layer \n    and the Thresholded Linear Unit (TLU).</p>\n\n<pre><code>Source: https://arxiv.org/pdf/1911.09737.pdf\n\n:param eps: An epsilon value to avoid division by zero\n:param weight_initializer: Initializer for the weights.\n:param weight_regularizer: Optional regularizer for the weights.\n:param weight_constraint: Optional constraint for the weights.\n:param bias_initializer: Initializer for the bias.\n:param bias_regularizer: Optional regularizer for the bias.\n:param bias_constraint: Optional constraint for the bias.\n:param threshold_initializer: Initializer for the beta threshold.\n:param threshold_regularizer: Optional regularizer for the threshold.\n:param threshold_constraint: Optional constraint for the threshold.\n\"\"\"\ndef __init__(self, \n             eps=1e-15, \n             weight_initializer='ones',\n             weight_regularizer=None,\n             weight_constraint=None,\n             bias_initializer='zeros',\n             bias_regularizer=None,\n             bias_constraint=None,\n             threshold_initializer='zeros',\n             threshold_regularizer=None,\n             threshold_constraint=None,\n             **kwargs):\n    super(FilterResponseNormalization, self).__init__(**kwargs)\n    self.supports_masking = True\n    self.eps = K.variable(eps, dtype=K.floatx())\n    self.weight_initializer = initializers.get(weight_initializer)\n    self.weight_regularizer = regularizers.get(weight_regularizer)\n    self.weight_constraint = constraints.get(weight_constraint)\n    self.bias_initializer = initializers.get(bias_initializer)\n    self.bias_regularizer = regularizers.get(bias_regularizer)\n    self.bias_constraint = constraints.get(bias_constraint)\n    self.threshold_constraint = constraints.get(threshold_constraint)\n    self.threshold_regularizer = regularizers.get(threshold_regularizer)\n    self.threshold_initializer = initializers.get(threshold_initializer)\n\ndef build(self, input_shape):\n    \"\"\"\n    Intialize weights, bias and threshold variables\n\n    :param input_shape: The shape of the input that this layer takes\n    \"\"\"\n    shape = self.input_shape[-1:]\n    self.weights = self.add_weight(shape=shape,\n                                   initializer=self.weight_initializer,\n                                   regularizer=self.weight_regularizer,\n                                   constraint=self.weight_constraint,\n                                   name='weights')\n    self.bias = self.add_weight(shape=shape,\n                                initializer=self.bias_initializer,\n                                regularizer=self.bias_regularizer,\n                                constraint=self.bias_constraint,\n                                name='bias')\n    self.threshold = self.add_weight(shape=shape,\n                               initializer=self.threshold_initializer,\n                               regularizer=self.threshold_regularizer,\n                               constraint=self.threshold_constraint,\n                               name='threshold')\n    super(LayerNormalization, self).build(input_shape)\n\ndef call(self, x):\n    \"\"\"\n    :param x: Input tensor of shape [NxHxWxC]\n    :return: the Filter Response Normalization with Thresholded Linear Unit activation\n    \"\"\"\n    # Compute the mean norm of activations per channel.\n    nu2 = K.reduce_mean(K.square(x), axis=[1, 2], keepdims=True)\n    # Perform FRN\n    x = x * K.rsqrt(nu2 + K.abs(eps))\n    # Perform TLU activation\n    x = self._tlu(x, weights=self.weights, biases=self.bias, threshold=self.threshold)\n    return x\n\ndef get_config(self):\n    config = {\n        'epsilon': self.epsilon,\n        'beta_initializer': initializers.serialize(self.bias_initializer),\n        'gamma_initializer': initializers.serialize(self.weight_initializer),\n        'tau_initializer': initializers.serialize(self.threshold_initializer),\n        'beta_regularizer': regularizers.serialize(self.bias_regularizer),\n        'gamma_regularizer': regularizers.serialize(self.weight_regularizer),\n        'tau_regularizer': regularizers.serialize(self.threshold_regularizer),\n        'beta_constraint': constraints.serialize(self.bias_constraint),\n        'gamma_constraint': constraints.serialize(self.weight_constraint),\n        'tau_constraint': constraints.serialize(self.threshold_constraint)\n    }\n    base_config = super(FilterResponseNormalization, self).get_config()\n    return dict(list(base_config.items()) + list(config.items()))\n\n@staticmethod\ndef compute_output_shape(input_shape):\n    \"\"\"\n    The output shape will be the same as the input shape.\n    \"\"\"\n    return input_shape\n\n@staticmethod\ndef _tlu(x, weights, bias, threshold):\n    \"\"\"\n    The Thresholded Linear Unit activation (TLU)\n    Source: https://arxiv.org/pdf/1911.09737.pdf\n\n    :param x: The input transformed by the Filtered Response Normalization\n    :param weights: The current weights of the layer\n    :param bias: The current biases of the layer\n    :param threshold: A learned threshold\n    :return: The output of the layer\n    \"\"\"\n    return K.maximum(weights * x + bias, threshold)\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Quick code snippet:\n\n```\nfrom keras import backend as K\nfrom keras.engine import Layer\nfrom keras import initializers, regularizers, constraints\n\n\nclass FilterResponseNormalization(Layer):\n    \"\"\"\n    Implementation of the Filter Response Normalization (FRN) layer \n    and the Thresholded Linear Unit (TLU).\n    \n    Source: https://arxiv.org/pdf/1911.09737.pdf\n    \n    :param eps: An epsilon value to avoid division by zero\n    :param weight_initializer: Initializer for the weights.\n    :param weight_regularizer: Optional regularizer for the weights.\n    :param weight_constraint: Optional constraint for the weights.\n    :param bias_initializer: Initializer for the bias.\n    :param bias_regularizer: Optional regularizer for the bias.\n    :param bias_constraint: Optional constraint for the bias.\n    :param threshold_initializer: Initializer for the beta threshold.\n    :param threshold_regularizer: Optional regularizer for the threshold.\n    :param threshold_constraint: Optional constraint for the threshold.\n    \"\"\"\n    def __init__(self, \n                 eps=1e-15, \n                 weight_initializer='ones',\n                 weight_regularizer=None,\n                 weight_constraint=None,\n                 bias_initializer='zeros',\n                 bias_regularizer=None,\n                 bias_constraint=None,\n                 threshold_initializer='zeros',\n                 threshold_regularizer=None,\n                 threshold_constraint=None,\n                 **kwargs):\n        super(FilterResponseNormalization, self).__init__(**kwargs)\n        self.supports_masking = True\n        self.eps = K.variable(eps, dtype=K.floatx())\n        self.weight_initializer = initializers.get(weight_initializer)\n        self.weight_regularizer = regularizers.get(weight_regularizer)\n        self.weight_constraint = constraints.get(weight_constraint)\n        self.bias_initializer = initializers.get(bias_initializer)\n        self.bias_regularizer = regularizers.get(bias_regularizer)\n        self.bias_constraint = constraints.get(bias_constraint)\n        self.threshold_constraint = constraints.get(threshold_constraint)\n        self.threshold_regularizer = regularizers.get(threshold_regularizer)\n        self.threshold_initializer = initializers.get(threshold_initializer)\n        \n    def build(self, input_shape):\n        \"\"\"\n        Intialize weights, bias and threshold variables\n        \n        :param input_shape: The shape of the input that this layer takes\n        \"\"\"\n        shape = self.input_shape[-1:]\n        self.weights = self.add_weight(shape=shape,\n                                       initializer=self.weight_initializer,\n                                       regularizer=self.weight_regularizer,\n                                       constraint=self.weight_constraint,\n                                       name='weights')\n        self.bias = self.add_weight(shape=shape,\n                                    initializer=self.bias_initializer,\n                                    regularizer=self.bias_regularizer,\n                                    constraint=self.bias_constraint,\n                                    name='bias')\n        self.threshold = self.add_weight(shape=shape,\n                                   initializer=self.threshold_initializer,\n                                   regularizer=self.threshold_regularizer,\n                                   constraint=self.threshold_constraint,\n                                   name='threshold')\n        super(LayerNormalization, self).build(input_shape)\n        \n    def call(self, x):\n        \"\"\"\n        :param x: Input tensor of shape [NxHxWxC]\n        :return: the Filter Response Normalization with Thresholded Linear Unit activation\n        \"\"\"\n        # Compute the mean norm of activations per channel.\n        nu2 = K.reduce_mean(K.square(x), axis=[1, 2], keepdims=True)\n        # Perform FRN\n        x = x * K.rsqrt(nu2 + K.abs(eps))\n        # Perform TLU activation\n        x = self._tlu(x, weights=self.weights, biases=self.bias, threshold=self.threshold)\n        return x\n    \n    def get_config(self):\n        config = {\n            'epsilon': self.epsilon,\n            'beta_initializer': initializers.serialize(self.bias_initializer),\n            'gamma_initializer': initializers.serialize(self.weight_initializer),\n            'tau_initializer': initializers.serialize(self.threshold_initializer),\n            'beta_regularizer': regularizers.serialize(self.bias_regularizer),\n            'gamma_regularizer': regularizers.serialize(self.weight_regularizer),\n            'tau_regularizer': regularizers.serialize(self.threshold_regularizer),\n            'beta_constraint': constraints.serialize(self.bias_constraint),\n            'gamma_constraint': constraints.serialize(self.weight_constraint),\n            'tau_constraint': constraints.serialize(self.threshold_constraint)\n        }\n        base_config = super(FilterResponseNormalization, self).get_config()\n        return dict(list(base_config.items()) + list(config.items()))\n        \n    @staticmethod\n    def compute_output_shape(input_shape):\n        \"\"\"\n        The output shape will be the same as the input shape.\n        \"\"\"\n        return input_shape\n    \n    @staticmethod\n    def _tlu(x, weights, bias, threshold):\n        \"\"\"\n        The Thresholded Linear Unit activation (TLU)\n        Source: https://arxiv.org/pdf/1911.09737.pdf\n        \n        :param x: The input transformed by the Filtered Response Normalization\n        :param weights: The current weights of the layer\n        :param bias: The current biases of the layer\n        :param threshold: A learned threshold\n        :return: The output of the layer\n        \"\"\"\n        return K.maximum(weights * x + bias, threshold)\n```",
      "votes": null
    },
    {
      "id": "688497",
      "postDate": "12/05/2019 16:30:58",
      "content": "<p>wow!!!!!!!!!!!!\npytorch implementation\ndid you try to use it using pytorch?\ni have learnt a good tactic from you from you kernel of aptos competition,you used GN there by replacing batch norm\nare you doing the same in this competition?\nif yes,can you tell me how good or bad it is doing here in this competition?thanks mate</p>",
      "rawMarkdown": "wow!!!!!!!!!!!!\npytorch implementation\ndid you try to use it using pytorch?\ni have learnt a good tactic from you from you kernel of aptos competition,you used GN there by replacing batch norm\nare you doing the same in this competition?\nif yes,can you tell me how good or bad it is doing here in this competition?thanks mate",
      "votes": null
    },
    {
      "id": "688530",
      "postDate": "12/05/2019 17:39:39",
      "content": "<p>👍 There are several PyTorch implementations on Github so you can easily find one, but I haven't tested it using PyTorch yet. </p>\n\n<p>Glad to hear that the APTOS kernel helped you! For the FRN layer I experimented on APTOS 2019. Using FRN with TLU instead of GN gave worse results. However, using only FRN + the original Swish activations in EfficientNetB5 gave me a slightly better private score, while it gave a slightly worse local validation score. </p>\n\n<p>TL;DR: FRN seems to generalize better than BN, GN, LN, etc. but TLU doesn't outperform Swish activations. </p>\n\n<p>To be honest, I'm not competing in the Baidu competition yet. It just seemed like the most appropriate place to post this. </p>\n\n<p>I'm happy to team up for this competition if you are interested. 🙂 </p>",
      "rawMarkdown": "👍 There are several PyTorch implementations on Github so you can easily find one, but I haven't tested it using PyTorch yet. \n\nGlad to hear that the APTOS kernel helped you! For the FRN layer I experimented on APTOS 2019. Using FRN with TLU instead of GN gave worse results. However, using only FRN + the original Swish activations in EfficientNetB5 gave me a slightly better private score, while it gave a slightly worse local validation score. \n\nTL;DR: FRN seems to generalize better than BN, GN, LN, etc. but TLU doesn't outperform Swish activations. \n\nTo be honest, I'm not competing in the Baidu competition yet. It just seemed like the most appropriate place to post this. \n\nI'm happy to team up for this competition if you are interested. 🙂",
      "votes": null
    },
    {
      "id": "688755",
      "postDate": "12/06/2019 03:14:29",
      "content": "<p>Thanks, I learn some new knowledge.😃 </p>",
      "rawMarkdown": "Thanks, I learn some new knowledge.😃",
      "votes": null
    },
    {
      "id": "688783",
      "postDate": "12/06/2019 05:04:42",
      "content": "<p>Wow. Thanks for sharing it <a href=\"/carlolepelaars\">@carlolepelaars</a>. </p>",
      "rawMarkdown": "Wow. Thanks for sharing it @carlolepelaars.",
      "votes": null
    },
    {
      "id": "689029",
      "postDate": "12/06/2019 11:22:45",
      "content": "<p>👍 </p>",
      "rawMarkdown": "👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 688461,
      "author_name": "carlolepelaars",
      "author_url": "",
      "post_date": "12/05/2019 16:02:34",
      "content": "<p>Quick code snippet:</p>\n\n<p>```\nfrom keras import backend as K\nfrom keras.engine import Layer\nfrom keras import initializers, regularizers, constraints</p>\n\n<p>class FilterResponseNormalization(Layer):\n    \"\"\"\n    Implementation of the Filter Response Normalization (FRN) layer \n    and the Thresholded Linear Unit (TLU).</p>\n\n<pre><code>Source: https://arxiv.org/pdf/1911.09737.pdf\n\n:param eps: An epsilon value to avoid division by zero\n:param weight_initializer: Initializer for the weights.\n:param weight_regularizer: Optional regularizer for the weights.\n:param weight_constraint: Optional constraint for the weights.\n:param bias_initializer: Initializer for the bias.\n:param bias_regularizer: Optional regularizer for the bias.\n:param bias_constraint: Optional constraint for the bias.\n:param threshold_initializer: Initializer for the beta threshold.\n:param threshold_regularizer: Optional regularizer for the threshold.\n:param threshold_constraint: Optional constraint for the threshold.\n\"\"\"\ndef __init__(self, \n             eps=1e-15, \n             weight_initializer='ones',\n             weight_regularizer=None,\n             weight_constraint=None,\n             bias_initializer='zeros',\n             bias_regularizer=None,\n             bias_constraint=None,\n             threshold_initializer='zeros',\n             threshold_regularizer=None,\n             threshold_constraint=None,\n             **kwargs):\n    super(FilterResponseNormalization, self).__init__(**kwargs)\n    self.supports_masking = True\n    self.eps = K.variable(eps, dtype=K.floatx())\n    self.weight_initializer = initializers.get(weight_initializer)\n    self.weight_regularizer = regularizers.get(weight_regularizer)\n    self.weight_constraint = constraints.get(weight_constraint)\n    self.bias_initializer = initializers.get(bias_initializer)\n    self.bias_regularizer = regularizers.get(bias_regularizer)\n    self.bias_constraint = constraints.get(bias_constraint)\n    self.threshold_constraint = constraints.get(threshold_constraint)\n    self.threshold_regularizer = regularizers.get(threshold_regularizer)\n    self.threshold_initializer = initializers.get(threshold_initializer)\n\ndef build(self, input_shape):\n    \"\"\"\n    Intialize weights, bias and threshold variables\n\n    :param input_shape: The shape of the input that this layer takes\n    \"\"\"\n    shape = self.input_shape[-1:]\n    self.weights = self.add_weight(shape=shape,\n                                   initializer=self.weight_initializer,\n                                   regularizer=self.weight_regularizer,\n                                   constraint=self.weight_constraint,\n                                   name='weights')\n    self.bias = self.add_weight(shape=shape,\n                                initializer=self.bias_initializer,\n                                regularizer=self.bias_regularizer,\n                                constraint=self.bias_constraint,\n                                name='bias')\n    self.threshold = self.add_weight(shape=shape,\n                               initializer=self.threshold_initializer,\n                               regularizer=self.threshold_regularizer,\n                               constraint=self.threshold_constraint,\n                               name='threshold')\n    super(LayerNormalization, self).build(input_shape)\n\ndef call(self, x):\n    \"\"\"\n    :param x: Input tensor of shape [NxHxWxC]\n    :return: the Filter Response Normalization with Thresholded Linear Unit activation\n    \"\"\"\n    # Compute the mean norm of activations per channel.\n    nu2 = K.reduce_mean(K.square(x), axis=[1, 2], keepdims=True)\n    # Perform FRN\n    x = x * K.rsqrt(nu2 + K.abs(eps))\n    # Perform TLU activation\n    x = self._tlu(x, weights=self.weights, biases=self.bias, threshold=self.threshold)\n    return x\n\ndef get_config(self):\n    config = {\n        'epsilon': self.epsilon,\n        'beta_initializer': initializers.serialize(self.bias_initializer),\n        'gamma_initializer': initializers.serialize(self.weight_initializer),\n        'tau_initializer': initializers.serialize(self.threshold_initializer),\n        'beta_regularizer': regularizers.serialize(self.bias_regularizer),\n        'gamma_regularizer': regularizers.serialize(self.weight_regularizer),\n        'tau_regularizer': regularizers.serialize(self.threshold_regularizer),\n        'beta_constraint': constraints.serialize(self.bias_constraint),\n        'gamma_constraint': constraints.serialize(self.weight_constraint),\n        'tau_constraint': constraints.serialize(self.threshold_constraint)\n    }\n    base_config = super(FilterResponseNormalization, self).get_config()\n    return dict(list(base_config.items()) + list(config.items()))\n\n@staticmethod\ndef compute_output_shape(input_shape):\n    \"\"\"\n    The output shape will be the same as the input shape.\n    \"\"\"\n    return input_shape\n\n@staticmethod\ndef _tlu(x, weights, bias, threshold):\n    \"\"\"\n    The Thresholded Linear Unit activation (TLU)\n    Source: https://arxiv.org/pdf/1911.09737.pdf\n\n    :param x: The input transformed by the Filtered Response Normalization\n    :param weights: The current weights of the layer\n    :param bias: The current biases of the layer\n    :param threshold: A learned threshold\n    :return: The output of the layer\n    \"\"\"\n    return K.maximum(weights * x + bias, threshold)\n</code></pre>\n\n<p>```</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 688497,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "12/05/2019 16:30:58",
      "content": "<p>wow!!!!!!!!!!!!\npytorch implementation\ndid you try to use it using pytorch?\ni have learnt a good tactic from you from you kernel of aptos competition,you used GN there by replacing batch norm\nare you doing the same in this competition?\nif yes,can you tell me how good or bad it is doing here in this competition?thanks mate</p>",
      "votes": null,
      "replies": [
        {
          "id": 688530,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "12/05/2019 17:39:39",
          "content": "<p>👍 There are several PyTorch implementations on Github so you can easily find one, but I haven't tested it using PyTorch yet. </p>\n\n<p>Glad to hear that the APTOS kernel helped you! For the FRN layer I experimented on APTOS 2019. Using FRN with TLU instead of GN gave worse results. However, using only FRN + the original Swish activations in EfficientNetB5 gave me a slightly better private score, while it gave a slightly worse local validation score. </p>\n\n<p>TL;DR: FRN seems to generalize better than BN, GN, LN, etc. but TLU doesn't outperform Swish activations. </p>\n\n<p>To be honest, I'm not competing in the Baidu competition yet. It just seemed like the most appropriate place to post this. </p>\n\n<p>I'm happy to team up for this competition if you are interested. 🙂 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 688755,
      "author_name": "diegojohnson",
      "author_url": "",
      "post_date": "12/06/2019 03:14:29",
      "content": "<p>Thanks, I learn some new knowledge.😃 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 688783,
      "author_name": "manojprabhaakr",
      "author_url": "",
      "post_date": "12/06/2019 05:04:42",
      "content": "<p>Wow. Thanks for sharing it <a href=\"/carlolepelaars\">@carlolepelaars</a>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 689029,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "12/06/2019 11:22:45",
          "content": "<p>👍 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "688459": "Last week Google AI released research on a new normalization layer called[ \"Filter Response Normalization\" (FRN) together with the \"Thresholded Linear Unit\" (TLU) activation](https://arxiv.org/pdf/1911.09737.pdf). They observed considerable gains in accuracy across all batch size, but I wanted to see for myself if it would work well on other data science problems. If the claims are true then it would make Batch Normalization, Group Normalization and Layer Normalization kind of obsolete.\n\nFRN\n\n![](https://deeplearn.org/arxiv_files/1911.09737v1/x1.png)\n\nImage: A comparison of FRN with BN, BRN and GN. [Source](https://deeplearn.org/arxiv/104635/filter-response-normalization-layer:-eliminating-batch-dependence-in-the-training-of-deep-neural-networks)\n\nI'm still experimenting with it, but just uploaded [an implementation of this on Github](https://github.com/CarloLepelaars/filter_response_normalization_keras). I tested it on EfficientNetB5 with a small batch size (4) and it seems to work well. However, I still have to do some more experiments to see if it is beneficial compared to group normalization or layer normalization with small batch sizes.\n\nPlease let me know here what your experiences are and if you see gains in performance from replacing your batch normalization and activation layers with FRN + TLU. Curious to see how well this research works in the real world!\n\nPaper:\nhttps://arxiv.org/pdf/1911.09737.pdf\n\nRepository with implementation using Tensorflow/Keras:\nhttps://github.com/CarloLepelaars/filter_response_normalization_keras",
    "688461": "Quick code snippet:\n\n```\nfrom keras import backend as K\nfrom keras.engine import Layer\nfrom keras import initializers, regularizers, constraints\n\n\nclass FilterResponseNormalization(Layer):\n    \"\"\"\n    Implementation of the Filter Response Normalization (FRN) layer \n    and the Thresholded Linear Unit (TLU).\n    \n    Source: https://arxiv.org/pdf/1911.09737.pdf\n    \n    :param eps: An epsilon value to avoid division by zero\n    :param weight_initializer: Initializer for the weights.\n    :param weight_regularizer: Optional regularizer for the weights.\n    :param weight_constraint: Optional constraint for the weights.\n    :param bias_initializer: Initializer for the bias.\n    :param bias_regularizer: Optional regularizer for the bias.\n    :param bias_constraint: Optional constraint for the bias.\n    :param threshold_initializer: Initializer for the beta threshold.\n    :param threshold_regularizer: Optional regularizer for the threshold.\n    :param threshold_constraint: Optional constraint for the threshold.\n    \"\"\"\n    def __init__(self, \n                 eps=1e-15, \n                 weight_initializer='ones',\n                 weight_regularizer=None,\n                 weight_constraint=None,\n                 bias_initializer='zeros',\n                 bias_regularizer=None,\n                 bias_constraint=None,\n                 threshold_initializer='zeros',\n                 threshold_regularizer=None,\n                 threshold_constraint=None,\n                 **kwargs):\n        super(FilterResponseNormalization, self).__init__(**kwargs)\n        self.supports_masking = True\n        self.eps = K.variable(eps, dtype=K.floatx())\n        self.weight_initializer = initializers.get(weight_initializer)\n        self.weight_regularizer = regularizers.get(weight_regularizer)\n        self.weight_constraint = constraints.get(weight_constraint)\n        self.bias_initializer = initializers.get(bias_initializer)\n        self.bias_regularizer = regularizers.get(bias_regularizer)\n        self.bias_constraint = constraints.get(bias_constraint)\n        self.threshold_constraint = constraints.get(threshold_constraint)\n        self.threshold_regularizer = regularizers.get(threshold_regularizer)\n        self.threshold_initializer = initializers.get(threshold_initializer)\n        \n    def build(self, input_shape):\n        \"\"\"\n        Intialize weights, bias and threshold variables\n        \n        :param input_shape: The shape of the input that this layer takes\n        \"\"\"\n        shape = self.input_shape[-1:]\n        self.weights = self.add_weight(shape=shape,\n                                       initializer=self.weight_initializer,\n                                       regularizer=self.weight_regularizer,\n                                       constraint=self.weight_constraint,\n                                       name='weights')\n        self.bias = self.add_weight(shape=shape,\n                                    initializer=self.bias_initializer,\n                                    regularizer=self.bias_regularizer,\n                                    constraint=self.bias_constraint,\n                                    name='bias')\n        self.threshold = self.add_weight(shape=shape,\n                                   initializer=self.threshold_initializer,\n                                   regularizer=self.threshold_regularizer,\n                                   constraint=self.threshold_constraint,\n                                   name='threshold')\n        super(LayerNormalization, self).build(input_shape)\n        \n    def call(self, x):\n        \"\"\"\n        :param x: Input tensor of shape [NxHxWxC]\n        :return: the Filter Response Normalization with Thresholded Linear Unit activation\n        \"\"\"\n        # Compute the mean norm of activations per channel.\n        nu2 = K.reduce_mean(K.square(x), axis=[1, 2], keepdims=True)\n        # Perform FRN\n        x = x * K.rsqrt(nu2 + K.abs(eps))\n        # Perform TLU activation\n        x = self._tlu(x, weights=self.weights, biases=self.bias, threshold=self.threshold)\n        return x\n    \n    def get_config(self):\n        config = {\n            'epsilon': self.epsilon,\n            'beta_initializer': initializers.serialize(self.bias_initializer),\n            'gamma_initializer': initializers.serialize(self.weight_initializer),\n            'tau_initializer': initializers.serialize(self.threshold_initializer),\n            'beta_regularizer': regularizers.serialize(self.bias_regularizer),\n            'gamma_regularizer': regularizers.serialize(self.weight_regularizer),\n            'tau_regularizer': regularizers.serialize(self.threshold_regularizer),\n            'beta_constraint': constraints.serialize(self.bias_constraint),\n            'gamma_constraint': constraints.serialize(self.weight_constraint),\n            'tau_constraint': constraints.serialize(self.threshold_constraint)\n        }\n        base_config = super(FilterResponseNormalization, self).get_config()\n        return dict(list(base_config.items()) + list(config.items()))\n        \n    @staticmethod\n    def compute_output_shape(input_shape):\n        \"\"\"\n        The output shape will be the same as the input shape.\n        \"\"\"\n        return input_shape\n    \n    @staticmethod\n    def _tlu(x, weights, bias, threshold):\n        \"\"\"\n        The Thresholded Linear Unit activation (TLU)\n        Source: https://arxiv.org/pdf/1911.09737.pdf\n        \n        :param x: The input transformed by the Filtered Response Normalization\n        :param weights: The current weights of the layer\n        :param bias: The current biases of the layer\n        :param threshold: A learned threshold\n        :return: The output of the layer\n        \"\"\"\n        return K.maximum(weights * x + bias, threshold)\n```",
    "688497": "wow!!!!!!!!!!!!\npytorch implementation\ndid you try to use it using pytorch?\ni have learnt a good tactic from you from you kernel of aptos competition,you used GN there by replacing batch norm\nare you doing the same in this competition?\nif yes,can you tell me how good or bad it is doing here in this competition?thanks mate",
    "688530": "👍 There are several PyTorch implementations on Github so you can easily find one, but I haven't tested it using PyTorch yet. \n\nGlad to hear that the APTOS kernel helped you! For the FRN layer I experimented on APTOS 2019. Using FRN with TLU instead of GN gave worse results. However, using only FRN + the original Swish activations in EfficientNetB5 gave me a slightly better private score, while it gave a slightly worse local validation score. \n\nTL;DR: FRN seems to generalize better than BN, GN, LN, etc. but TLU doesn't outperform Swish activations. \n\nTo be honest, I'm not competing in the Baidu competition yet. It just seemed like the most appropriate place to post this. \n\nI'm happy to team up for this competition if you are interested. 🙂",
    "688755": "Thanks, I learn some new knowledge.😃",
    "688783": "Wow. Thanks for sharing it @carlolepelaars.",
    "689029": "👍"
  },
  "source": "meta"
}