{
  "id": 317746,
  "title": "SmeLU: Real World Large Scale Recommendation Systems Reproducibility and Smooth Activations",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/317746",
  "author_name": "Sinan Calisir",
  "post_date": "2022-04-08T15:01:48.195000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi everyone, </p>\n<p>I would like to share a recent blog post that Google has published. In this blog post, the authors discuss the idea of reproducibility of the results (in recommender systems particularly) and propose a new activation function: Smooth reLU (SmeLU) that shows a boost in the accuracy while addressing the reproducibility problems.</p>\n<p><strong>Source:</strong> <a href=\"https://ai.googleblog.com/2022/04/reproducibility-in-deep-learning-and.html\" target=\"_blank\">https://ai.googleblog.com/2022/04/reproducibility-in-deep-learning-and.html</a><br>\n<strong>Arxiv:</strong> <a href=\"https://arxiv.org/abs/2202.06499\" target=\"_blank\">https://arxiv.org/abs/2202.06499</a></p>\n<p><strong>From the blog post:</strong></p>\n<p>In <a href=\"https://arxiv.org/abs/2202.06499\" target=\"_blank\">“Real World Large Scale Recommendation Systems Reproducibility and Smooth Activations”</a>, we consider a different practical solution to this problem that does not incur the costs of other solutions, while still improving reproducibility and yielding higher model accuracy. We discover that the <a href=\"https://en.wikipedia.org/wiki/Rectifier_(neural_networks)\" target=\"_blank\">Rectified Linear Unit (ReLU)</a>, which is very popular as the nonlinearity function (i.e., <a href=\"https://en.wikipedia.org/wiki/Activation_function\" target=\"_blank\">activation function</a>) used to transform values in neural networks, exacerbates the irreproducibility problem. On the other hand, we demonstrate that smooth activation functions, which have derivatives that are continuous for the whole domain, unlike those of ReLU, are able to substantially reduce irreproducibility levels. We then propose the Smooth reLU (SmeLU) activation function, which gives comparable reproducibility and accuracy benefits to other smooth activations but is much simpler.</p>\n<p>ReLU, which is not a smooth function, imposes an objective whose landscape is partitioned into many regions with multiple local minima, each providing different model predictions. With this landscape, the order in which updates are applied is a dominant factor in determining the optimization trajectory, providing a recipe for irreproducibility. Because of its non-continuous gradient, functions expressed by a ReLU network will contain sudden jumps in the gradient, which can occur internally in different layers of the deep network, affecting updates of different internal units, and are likely strong contributors to irreproducibility.</p>\n<p><img src=\"https://raw.githubusercontent.com/snnclsr/kaggle_images/main/smooth_activations.png\" alt=\"\"></p>\n<h4>Smooth reLU (SmeLU)</h4>\n<p>Activations like GELU and Swish require complex hardware implementations to support exponential and logarithmic functions. Further, GELU must be computed numerically or approximated. These properties can make deployment error-prone, expensive, or slow. GELU and Swish are not monotonic (they start by slightly decreasing and then switch to increasing), which may interfere with interpretability (or identifiability), nor do they have a full stop or a clean slope 1 region, properties that simplify implementation and may aid in reproducibility. </p>\n<p>The Smooth reLU (SmeLU) activation function is designed as a simple function that addresses the concerns with other smooth activations. It connects a 0 slope on the left with a slope 1 line on the right through a quadratic middle region, constraining continuous gradients at the connection points (as an asymmetric version of a Huber loss function).</p>\n<h4>Performance</h4>\n<p>SmeLU has benefited multiple systems, specifically recommendation systems, increasing their reproducibility by reducing, for example, recommendation swap rates. While the use of SmeLU results in accuracy improvements over ReLU, it also replaces other costly methods to address irreproducibility, such as ensembles, which mitigate irreproducibility at the cost of accuracy. Moreover, replacing ensembles in sparse recommendation systems reduces the need for multiple lookups of model parameters that are needed to generate an inference for each of the ensemble components. This substantially improves training and inference efficiency.</p>",
  "messages": [
    {
      "id": 1749423,
      "postDate": "2022-04-08T15:01:48.197Z",
      "content": "<p>Hi everyone, </p>\n<p>I would like to share a recent blog post that Google has published. In this blog post, the authors discuss the idea of reproducibility of the results (in recommender systems particularly) and propose a new activation function: Smooth reLU (SmeLU) that shows a boost in the accuracy while addressing the reproducibility problems.</p>\n<p><strong>Source:</strong> <a href=\"https://ai.googleblog.com/2022/04/reproducibility-in-deep-learning-and.html\" target=\"_blank\">https://ai.googleblog.com/2022/04/reproducibility-in-deep-learning-and.html</a><br>\n<strong>Arxiv:</strong> <a href=\"https://arxiv.org/abs/2202.06499\" target=\"_blank\">https://arxiv.org/abs/2202.06499</a></p>\n<p><strong>From the blog post:</strong></p>\n<p>In <a href=\"https://arxiv.org/abs/2202.06499\" target=\"_blank\">“Real World Large Scale Recommendation Systems Reproducibility and Smooth Activations”</a>, we consider a different practical solution to this problem that does not incur the costs of other solutions, while still improving reproducibility and yielding higher model accuracy. We discover that the <a href=\"https://en.wikipedia.org/wiki/Rectifier_(neural_networks)\" target=\"_blank\">Rectified Linear Unit (ReLU)</a>, which is very popular as the nonlinearity function (i.e., <a href=\"https://en.wikipedia.org/wiki/Activation_function\" target=\"_blank\">activation function</a>) used to transform values in neural networks, exacerbates the irreproducibility problem. On the other hand, we demonstrate that smooth activation functions, which have derivatives that are continuous for the whole domain, unlike those of ReLU, are able to substantially reduce irreproducibility levels. We then propose the Smooth reLU (SmeLU) activation function, which gives comparable reproducibility and accuracy benefits to other smooth activations but is much simpler.</p>\n<p>ReLU, which is not a smooth function, imposes an objective whose landscape is partitioned into many regions with multiple local minima, each providing different model predictions. With this landscape, the order in which updates are applied is a dominant factor in determining the optimization trajectory, providing a recipe for irreproducibility. Because of its non-continuous gradient, functions expressed by a ReLU network will contain sudden jumps in the gradient, which can occur internally in different layers of the deep network, affecting updates of different internal units, and are likely strong contributors to irreproducibility.</p>\n<p><img src=\"https://raw.githubusercontent.com/snnclsr/kaggle_images/main/smooth_activations.png\" alt=\"\"></p>\n<h4>Smooth reLU (SmeLU)</h4>\n<p>Activations like GELU and Swish require complex hardware implementations to support exponential and logarithmic functions. Further, GELU must be computed numerically or approximated. These properties can make deployment error-prone, expensive, or slow. GELU and Swish are not monotonic (they start by slightly decreasing and then switch to increasing), which may interfere with interpretability (or identifiability), nor do they have a full stop or a clean slope 1 region, properties that simplify implementation and may aid in reproducibility. </p>\n<p>The Smooth reLU (SmeLU) activation function is designed as a simple function that addresses the concerns with other smooth activations. It connects a 0 slope on the left with a slope 1 line on the right through a quadratic middle region, constraining continuous gradients at the connection points (as an asymmetric version of a Huber loss function).</p>\n<h4>Performance</h4>\n<p>SmeLU has benefited multiple systems, specifically recommendation systems, increasing their reproducibility by reducing, for example, recommendation swap rates. While the use of SmeLU results in accuracy improvements over ReLU, it also replaces other costly methods to address irreproducibility, such as ensembles, which mitigate irreproducibility at the cost of accuracy. Moreover, replacing ensembles in sparse recommendation systems reduces the need for multiple lookups of model parameters that are needed to generate an inference for each of the ensemble components. This substantially improves training and inference efficiency.</p>",
      "rawMarkdown": "Hi everyone, \n\nI would like to share a recent blog post that Google has published. In this blog post, the authors discuss the idea of reproducibility of the results (in recommender systems particularly) and propose a new activation function: Smooth reLU (SmeLU) that shows a boost in the accuracy while addressing the reproducibility problems.\n\n**Source:** https://ai.googleblog.com/2022/04/reproducibility-in-deep-learning-and.html\n**Arxiv:** https://arxiv.org/abs/2202.06499\n\n\n**From the blog post:**\n\nIn [“Real World Large Scale Recommendation Systems Reproducibility and Smooth Activations”](https://arxiv.org/abs/2202.06499), we consider a different practical solution to this problem that does not incur the costs of other solutions, while still improving reproducibility and yielding higher model accuracy. We discover that the [Rectified Linear Unit (ReLU)](https://en.wikipedia.org/wiki/Rectifier_(neural_networks)), which is very popular as the nonlinearity function (i.e., [activation function](https://en.wikipedia.org/wiki/Activation_function)) used to transform values in neural networks, exacerbates the irreproducibility problem. On the other hand, we demonstrate that smooth activation functions, which have derivatives that are continuous for the whole domain, unlike those of ReLU, are able to substantially reduce irreproducibility levels. We then propose the Smooth reLU (SmeLU) activation function, which gives comparable reproducibility and accuracy benefits to other smooth activations but is much simpler.\n\nReLU, which is not a smooth function, imposes an objective whose landscape is partitioned into many regions with multiple local minima, each providing different model predictions. With this landscape, the order in which updates are applied is a dominant factor in determining the optimization trajectory, providing a recipe for irreproducibility. Because of its non-continuous gradient, functions expressed by a ReLU network will contain sudden jumps in the gradient, which can occur internally in different layers of the deep network, affecting updates of different internal units, and are likely strong contributors to irreproducibility.\n\n![](https://raw.githubusercontent.com/snnclsr/kaggle_images/main/smooth_activations.png)\n\n#### Smooth reLU (SmeLU)\nActivations like GELU and Swish require complex hardware implementations to support exponential and logarithmic functions. Further, GELU must be computed numerically or approximated. These properties can make deployment error-prone, expensive, or slow. GELU and Swish are not monotonic (they start by slightly decreasing and then switch to increasing), which may interfere with interpretability (or identifiability), nor do they have a full stop or a clean slope 1 region, properties that simplify implementation and may aid in reproducibility. \n\nThe Smooth reLU (SmeLU) activation function is designed as a simple function that addresses the concerns with other smooth activations. It connects a 0 slope on the left with a slope 1 line on the right through a quadratic middle region, constraining continuous gradients at the connection points (as an asymmetric version of a Huber loss function).\n\n#### Performance\n\nSmeLU has benefited multiple systems, specifically recommendation systems, increasing their reproducibility by reducing, for example, recommendation swap rates. While the use of SmeLU results in accuracy improvements over ReLU, it also replaces other costly methods to address irreproducibility, such as ensembles, which mitigate irreproducibility at the cost of accuracy. Moreover, replacing ensembles in sparse recommendation systems reduces the need for multiple lookups of model parameters that are needed to generate an inference for each of the ensemble components. This substantially improves training and inference efficiency.\n\n",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1749423": "Hi everyone, \n\nI would like to share a recent blog post that Google has published. In this blog post, the authors discuss the idea of reproducibility of the results (in recommender systems particularly) and propose a new activation function: Smooth reLU (SmeLU) that shows a boost in the accuracy while addressing the reproducibility problems.\n\n**Source:** https://ai.googleblog.com/2022/04/reproducibility-in-deep-learning-and.html\n**Arxiv:** https://arxiv.org/abs/2202.06499\n\n\n**From the blog post:**\n\nIn [“Real World Large Scale Recommendation Systems Reproducibility and Smooth Activations”](https://arxiv.org/abs/2202.06499), we consider a different practical solution to this problem that does not incur the costs of other solutions, while still improving reproducibility and yielding higher model accuracy. We discover that the [Rectified Linear Unit (ReLU)](https://en.wikipedia.org/wiki/Rectifier_(neural_networks)), which is very popular as the nonlinearity function (i.e., [activation function](https://en.wikipedia.org/wiki/Activation_function)) used to transform values in neural networks, exacerbates the irreproducibility problem. On the other hand, we demonstrate that smooth activation functions, which have derivatives that are continuous for the whole domain, unlike those of ReLU, are able to substantially reduce irreproducibility levels. We then propose the Smooth reLU (SmeLU) activation function, which gives comparable reproducibility and accuracy benefits to other smooth activations but is much simpler.\n\nReLU, which is not a smooth function, imposes an objective whose landscape is partitioned into many regions with multiple local minima, each providing different model predictions. With this landscape, the order in which updates are applied is a dominant factor in determining the optimization trajectory, providing a recipe for irreproducibility. Because of its non-continuous gradient, functions expressed by a ReLU network will contain sudden jumps in the gradient, which can occur internally in different layers of the deep network, affecting updates of different internal units, and are likely strong contributors to irreproducibility.\n\n![](https://raw.githubusercontent.com/snnclsr/kaggle_images/main/smooth_activations.png)\n\n#### Smooth reLU (SmeLU)\nActivations like GELU and Swish require complex hardware implementations to support exponential and logarithmic functions. Further, GELU must be computed numerically or approximated. These properties can make deployment error-prone, expensive, or slow. GELU and Swish are not monotonic (they start by slightly decreasing and then switch to increasing), which may interfere with interpretability (or identifiability), nor do they have a full stop or a clean slope 1 region, properties that simplify implementation and may aid in reproducibility. \n\nThe Smooth reLU (SmeLU) activation function is designed as a simple function that addresses the concerns with other smooth activations. It connects a 0 slope on the left with a slope 1 line on the right through a quadratic middle region, constraining continuous gradients at the connection points (as an asymmetric version of a Huber loss function).\n\n#### Performance\n\nSmeLU has benefited multiple systems, specifically recommendation systems, increasing their reproducibility by reducing, for example, recommendation swap rates. While the use of SmeLU results in accuracy improvements over ReLU, it also replaces other costly methods to address irreproducibility, such as ensembles, which mitigate irreproducibility at the cost of accuracy. Moreover, replacing ensembles in sparse recommendation systems reduces the need for multiple lookups of model parameters that are needed to generate an inference for each of the ensemble components. This substantially improves training and inference efficiency.\n\n"
  }
}