{
  "id": 205808,
  "title": "[Info]: Attention-based Rectified-Linear-Unit",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/205808",
  "author_name": "",
  "post_date": "2020-12-22T00:13:06.910068100Z",
  "votes": 24,
  "comment_count": 4,
  "views": 0,
  "content": "<p>A recent work, quite interesting. In case you're interested, check it out: <a href=\"https://arxiv.org/pdf/2006.13858.pdf\" target=\"_blank\"><strong>ARELU:</strong></a>, <a href=\"https://github.com/densechen/AReLU\" target=\"_blank\">code</a>. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fde889b3780cca3ee13c8505e6798b7e6%2F1.png?generation=1608595823213095&amp;alt=media\" alt=\"\"></p>\n<h3>Pytorch Implementation [Official]</h3>\n<pre><code>class AReLU(nn.Module):\n    def __init__(self, alpha=0.90, beta=2.0):\n        super().__init__()\n        self.alpha = nn.Parameter(torch.tensor([alpha]))\n        self.beta = nn.Parameter(torch.tensor([beta]))\n\n    def forward(self, input):\n        alpha = torch.clamp(self.alpha, min=0.01, max=0.99)\n        beta = 1 + torch.sigmoid(self.beta)\n\n        return F.relu(input) * beta - F.relu(-input) * alpha\n</code></pre>\n<h3>TensorFlow/Keras Implementation [Un-official]</h3>\n<pre><code># by a function \ndef ARelu(x, alpha=0.90, beta=2.0):\n    alpha = tf.clip_by_value(alpha, clip_value_min=0.01, \n                                           clip_value_max=0.99)\n    beta  = 1 + tf.math.sigmoid(beta)\n    return tf.nn.relu(x) * beta - tf.nn.relu(-x) * alpha\n</code></pre>\n<pre><code># or by a custom layer\nclass ARelu(tf.keras.layers.Layer):\n    def __init__(self, alpha=0.90, beta=2.0, **kwargs):\n        super(ARelu, self).__init__(**kwargs)\n        self.alpha = 0.90\n        self.beta  = 2.0\n\n    def call(self, inputs, training=None):\n        alpha = tf.clip_by_value(self.alpha, clip_value_min=0.01, \n                                               clip_value_max=0.99)\n        beta  = 1 + tf.math.sigmoid(self.beta)\n        return tf.nn.relu(inputs) * beta - tf.nn.relu(-inputs) * alpha\n</code></pre>",
  "messages": [
    {
      "id": "1121831",
      "postDate": "12/22/2020 00:13:06",
      "content": "<p>A recent work, quite interesting. In case you're interested, check it out: <a href=\"https://arxiv.org/pdf/2006.13858.pdf\" target=\"_blank\"><strong>ARELU:</strong></a>, <a href=\"https://github.com/densechen/AReLU\" target=\"_blank\">code</a>. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fde889b3780cca3ee13c8505e6798b7e6%2F1.png?generation=1608595823213095&amp;alt=media\" alt=\"\"></p>\n<h3>Pytorch Implementation [Official]</h3>\n<pre><code>class AReLU(nn.Module):\n    def __init__(self, alpha=0.90, beta=2.0):\n        super().__init__()\n        self.alpha = nn.Parameter(torch.tensor([alpha]))\n        self.beta = nn.Parameter(torch.tensor([beta]))\n\n    def forward(self, input):\n        alpha = torch.clamp(self.alpha, min=0.01, max=0.99)\n        beta = 1 + torch.sigmoid(self.beta)\n\n        return F.relu(input) * beta - F.relu(-input) * alpha\n</code></pre>\n<h3>TensorFlow/Keras Implementation [Un-official]</h3>\n<pre><code># by a function \ndef ARelu(x, alpha=0.90, beta=2.0):\n    alpha = tf.clip_by_value(alpha, clip_value_min=0.01, \n                                           clip_value_max=0.99)\n    beta  = 1 + tf.math.sigmoid(beta)\n    return tf.nn.relu(x) * beta - tf.nn.relu(-x) * alpha\n</code></pre>\n<pre><code># or by a custom layer\nclass ARelu(tf.keras.layers.Layer):\n    def __init__(self, alpha=0.90, beta=2.0, **kwargs):\n        super(ARelu, self).__init__(**kwargs)\n        self.alpha = 0.90\n        self.beta  = 2.0\n\n    def call(self, inputs, training=None):\n        alpha = tf.clip_by_value(self.alpha, clip_value_min=0.01, \n                                               clip_value_max=0.99)\n        beta  = 1 + tf.math.sigmoid(self.beta)\n        return tf.nn.relu(inputs) * beta - tf.nn.relu(-inputs) * alpha\n</code></pre>",
      "rawMarkdown": "A recent work, quite interesting. In case you're interested, check it out: [**ARELU:**](https://arxiv.org/pdf/2006.13858.pdf), [code](https://github.com/densechen/AReLU). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fde889b3780cca3ee13c8505e6798b7e6%2F1.png?generation=1608595823213095&alt=media)\n\n### Pytorch Implementation [Official]\n\n```\nclass AReLU(nn.Module):\n    def __init__(self, alpha=0.90, beta=2.0):\n        super().__init__()\n        self.alpha = nn.Parameter(torch.tensor([alpha]))\n        self.beta = nn.Parameter(torch.tensor([beta]))\n\n    def forward(self, input):\n        alpha = torch.clamp(self.alpha, min=0.01, max=0.99)\n        beta = 1 + torch.sigmoid(self.beta)\n\n        return F.relu(input) * beta - F.relu(-input) * alpha\n```\n\n### TensorFlow/Keras Implementation [Un-official]\n\n```\n# by a function \ndef ARelu(x, alpha=0.90, beta=2.0):\n    alpha = tf.clip_by_value(alpha, clip_value_min=0.01, \n                                           clip_value_max=0.99)\n    beta  = 1 + tf.math.sigmoid(beta)\n    return tf.nn.relu(x) * beta - tf.nn.relu(-x) * alpha\n```\n\n```\n# or by a custom layer\nclass ARelu(tf.keras.layers.Layer):\n    def __init__(self, alpha=0.90, beta=2.0, **kwargs):\n        super(ARelu, self).__init__(**kwargs)\n        self.alpha = 0.90\n        self.beta  = 2.0\n\n    def call(self, inputs, training=None):\n        alpha = tf.clip_by_value(self.alpha, clip_value_min=0.01, \n                                               clip_value_max=0.99)\n        beta  = 1 + tf.math.sigmoid(self.beta)\n        return tf.nn.relu(inputs) * beta - tf.nn.relu(-inputs) * alpha\n```",
      "votes": null
    },
    {
      "id": "1124340",
      "postDate": "12/23/2020 20:51:32",
      "content": "<p>Thank you for sharing, have to try it out.<br>\nFun fact : pytorch uses the name clamp and tensorflow uses the name clip. <br>\nclamp, clip 👍</p>",
      "rawMarkdown": "Thank you for sharing, have to try it out.\nFun fact : pytorch uses the name clamp and tensorflow uses the name clip. \nclamp, clip 👍",
      "votes": null
    },
    {
      "id": "1124983",
      "postDate": "12/24/2020 10:02:31",
      "content": "<p>Thanks for sharing. <br>\nTo readers who wonder what the function looks like: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1700650%2F5427cb8f8d7e08f57c4255b94e912816%2Farelu.png?generation=1608804119069252&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks for sharing. \nTo readers who wonder what the function looks like: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1700650%2F5427cb8f8d7e08f57c4255b94e912816%2Farelu.png?generation=1608804119069252&alt=media)",
      "votes": null
    },
    {
      "id": "1128278",
      "postDate": "12/27/2020 09:44:13",
      "content": "<p>Very good! What is it's key idea? Why they benefits come from?</p>",
      "rawMarkdown": "Very good! What is it's key idea? Why they benefits come from?",
      "votes": null
    },
    {
      "id": "1129934",
      "postDate": "12/28/2020 16:53:16",
      "content": "<p>From the abstraction</p>\n<blockquote>\n  <p>We propose a new perspective of <strong>learnable activation function</strong> through formulating<br>\n  them with <strong>element-wise</strong> attention mechanism. In each network layer, we devise<br>\n  an attention module which learns an element-wise, <strong>sign-based attention map</strong> for<br>\n  the pre-activation feature map. The attention map scales an element based on its<br>\n  sign. Adding the attention module with a rectified linear unit (ReLU) results in<br>\n  an <strong>amplification of positive elements and a suppression of negative ones</strong>, both<br>\n  with learned, data-adaptive parameters. </p>\n</blockquote>\n<p>You can read the <a href=\"https://arxiv.org/pdf/2006.13858.pdf\" target=\"_blank\">paper</a>, they provide comprehensive studies. </p>",
      "rawMarkdown": "From the abstraction\n\n> We propose a new perspective of **learnable activation function** through formulating\nthem with **element-wise** attention mechanism. In each network layer, we devise\nan attention module which learns an element-wise, **sign-based attention map** for\nthe pre-activation feature map. The attention map scales an element based on its\nsign. Adding the attention module with a rectified linear unit (ReLU) results in\nan **amplification of positive elements and a suppression of negative ones**, both\nwith learned, data-adaptive parameters. \n\nYou can read the [paper](https://arxiv.org/pdf/2006.13858.pdf), they provide comprehensive studies.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1124340,
      "author_name": "alincijov",
      "author_url": "",
      "post_date": "12/23/2020 20:51:32",
      "content": "<p>Thank you for sharing, have to try it out.<br>\nFun fact : pytorch uses the name clamp and tensorflow uses the name clip. <br>\nclamp, clip 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1124983,
      "author_name": "levelsofdescription",
      "author_url": "",
      "post_date": "12/24/2020 10:02:31",
      "content": "<p>Thanks for sharing. <br>\nTo readers who wonder what the function looks like: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1700650%2F5427cb8f8d7e08f57c4255b94e912816%2Farelu.png?generation=1608804119069252&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1128278,
      "author_name": "zavodrobotov",
      "author_url": "",
      "post_date": "12/27/2020 09:44:13",
      "content": "<p>Very good! What is it's key idea? Why they benefits come from?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1129934,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "12/28/2020 16:53:16",
          "content": "<p>From the abstraction</p>\n<blockquote>\n  <p>We propose a new perspective of <strong>learnable activation function</strong> through formulating<br>\n  them with <strong>element-wise</strong> attention mechanism. In each network layer, we devise<br>\n  an attention module which learns an element-wise, <strong>sign-based attention map</strong> for<br>\n  the pre-activation feature map. The attention map scales an element based on its<br>\n  sign. Adding the attention module with a rectified linear unit (ReLU) results in<br>\n  an <strong>amplification of positive elements and a suppression of negative ones</strong>, both<br>\n  with learned, data-adaptive parameters. </p>\n</blockquote>\n<p>You can read the <a href=\"https://arxiv.org/pdf/2006.13858.pdf\" target=\"_blank\">paper</a>, they provide comprehensive studies. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1121831": "A recent work, quite interesting. In case you're interested, check it out: [**ARELU:**](https://arxiv.org/pdf/2006.13858.pdf), [code](https://github.com/densechen/AReLU). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fde889b3780cca3ee13c8505e6798b7e6%2F1.png?generation=1608595823213095&alt=media)\n\n### Pytorch Implementation [Official]\n\n```\nclass AReLU(nn.Module):\n    def __init__(self, alpha=0.90, beta=2.0):\n        super().__init__()\n        self.alpha = nn.Parameter(torch.tensor([alpha]))\n        self.beta = nn.Parameter(torch.tensor([beta]))\n\n    def forward(self, input):\n        alpha = torch.clamp(self.alpha, min=0.01, max=0.99)\n        beta = 1 + torch.sigmoid(self.beta)\n\n        return F.relu(input) * beta - F.relu(-input) * alpha\n```\n\n### TensorFlow/Keras Implementation [Un-official]\n\n```\n# by a function \ndef ARelu(x, alpha=0.90, beta=2.0):\n    alpha = tf.clip_by_value(alpha, clip_value_min=0.01, \n                                           clip_value_max=0.99)\n    beta  = 1 + tf.math.sigmoid(beta)\n    return tf.nn.relu(x) * beta - tf.nn.relu(-x) * alpha\n```\n\n```\n# or by a custom layer\nclass ARelu(tf.keras.layers.Layer):\n    def __init__(self, alpha=0.90, beta=2.0, **kwargs):\n        super(ARelu, self).__init__(**kwargs)\n        self.alpha = 0.90\n        self.beta  = 2.0\n\n    def call(self, inputs, training=None):\n        alpha = tf.clip_by_value(self.alpha, clip_value_min=0.01, \n                                               clip_value_max=0.99)\n        beta  = 1 + tf.math.sigmoid(self.beta)\n        return tf.nn.relu(inputs) * beta - tf.nn.relu(-inputs) * alpha\n```",
    "1124340": "Thank you for sharing, have to try it out.\nFun fact : pytorch uses the name clamp and tensorflow uses the name clip. \nclamp, clip 👍",
    "1124983": "Thanks for sharing. \nTo readers who wonder what the function looks like: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1700650%2F5427cb8f8d7e08f57c4255b94e912816%2Farelu.png?generation=1608804119069252&alt=media)",
    "1128278": "Very good! What is it's key idea? Why they benefits come from?",
    "1129934": "From the abstraction\n\n> We propose a new perspective of **learnable activation function** through formulating\nthem with **element-wise** attention mechanism. In each network layer, we devise\nan attention module which learns an element-wise, **sign-based attention map** for\nthe pre-activation feature map. The attention map scales an element based on its\nsign. Adding the attention module with a rectified linear unit (ReLU) results in\nan **amplification of positive elements and a suppression of negative ones**, both\nwith learned, data-adaptive parameters. \n\nYou can read the [paper](https://arxiv.org/pdf/2006.13858.pdf), they provide comprehensive studies."
  },
  "source": "meta"
}