{
  "id": 360084,
  "title": "NN activation functions comparison",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/360084",
  "author_name": "",
  "post_date": "2022-10-14T22:13:10.450157300Z",
  "votes": 20,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi Kaggle fellows, </p>\n<p>In this thread I would like to discuss the differences in the NN activation functions, as currently we have <a href=\"https://www.kaggle.com/code/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta\" target=\"_blank\">the top-1 kernel</a> based on the NN built with FastAI framework.</p>\n<p>When you try to create the better NN model, you almost always try to find the best NN structure you can - and this is not only the number of layers and their sizes! It's also <strong>important to choose the right activation for the layers</strong> to prevent the usual NN problems like vanishing or exploding gradients, bad training speed or final quality.  </p>\n<p>What activations do we know, how do they look like in comparison to each other?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fac63fa4bb56c678383c0b68a163c5f4e%2Factivation.png?generation=1665783746972764&amp;alt=media\" alt=\"\"> </p>\n<p>The most commonly known variants are Sigmoid, Tanh and ReLU, but all of them have some advantages and disadvantages. Let's take a closer look for the ReLU, the most common default activation for the frameworks nowadays.</p>\n<p><strong>Advantages of ReLU:</strong></p>\n<ul>\n<li>For <code>x &gt; 0</code> ReLU has constant gradient, which reduces the chance of vanishing gradients at any point of time (results in faster learning)</li>\n<li>For <code>x &lt;= 0</code> the gradient is zero and hence less number of neurons will be fired which reduce overfitting and is also cost efficient</li>\n<li>ReLU has better convergence performance as compared to other activation functions</li>\n</ul>\n<p><strong>Disadvantage of ReLU</strong> - the dying ReLU problem. A ReLU neuron is dead if it is stuck in the negative side and always outputs zero (so it is impossible for the neuron to recover back)</p>\n<p><strong>To fix the ReLU disadvantage</strong> in 2017 Google Brain invented the <code>Swish</code> activation - almost <code>ReLU</code>, but with fixed negative part and the parameter to change the slope for the positives. It shows better scores on the different benchmarks like ImageNet and everybody was happy about that.</p>\n<p>But still researchers wanted <strong>better accuracy, stability and universality</strong> so in 2019 the <code>Mish</code> activation appears. Why it can be better than <code>Swish</code>?</p>\n<ul>\n<li><code>Mish</code> outperforms in case of noisy input conditions in comparison with <code>ReLU</code> and <code>Swish</code></li>\n<li>Test accuracy with different optimizers has less drop in case of <code>Mish</code> compared with <code>Swish</code></li>\n<li><code>Mish</code> has shown to have a consistent improvement over <code>Swish</code> using different dropout rates</li>\n<li>Test accuracy with different learning rates for <code>Mish</code> is also usually better than <code>Swish</code></li>\n</ul>\n<p>So that's why <code>Mish</code> become the better, more stable and universal activation for NNs but it is also important to say - <strong>to find the optimal activation function for the specific ML task/competition, you need to check different variants and compare them on your validation set</strong>. The data and the experiment should say what is the best choice, not the papers.</p>\n<p>Hope this helps you to make a better choice :)</p>\n<p>Alex</p>\n<p>P.S. For more details and graphs you can check <a href=\"https://krutikabapat.github.io/Swish-Vs-Mish-Latest-Activation-Functions/\" target=\"_blank\">this link</a></p>",
  "messages": [
    {
      "id": "1987904",
      "postDate": "10/14/2022 22:13:10",
      "content": "<p>Hi Kaggle fellows, </p>\n<p>In this thread I would like to discuss the differences in the NN activation functions, as currently we have <a href=\"https://www.kaggle.com/code/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta\" target=\"_blank\">the top-1 kernel</a> based on the NN built with FastAI framework.</p>\n<p>When you try to create the better NN model, you almost always try to find the best NN structure you can - and this is not only the number of layers and their sizes! It's also <strong>important to choose the right activation for the layers</strong> to prevent the usual NN problems like vanishing or exploding gradients, bad training speed or final quality.  </p>\n<p>What activations do we know, how do they look like in comparison to each other?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fac63fa4bb56c678383c0b68a163c5f4e%2Factivation.png?generation=1665783746972764&amp;alt=media\" alt=\"\"> </p>\n<p>The most commonly known variants are Sigmoid, Tanh and ReLU, but all of them have some advantages and disadvantages. Let's take a closer look for the ReLU, the most common default activation for the frameworks nowadays.</p>\n<p><strong>Advantages of ReLU:</strong></p>\n<ul>\n<li>For <code>x &gt; 0</code> ReLU has constant gradient, which reduces the chance of vanishing gradients at any point of time (results in faster learning)</li>\n<li>For <code>x &lt;= 0</code> the gradient is zero and hence less number of neurons will be fired which reduce overfitting and is also cost efficient</li>\n<li>ReLU has better convergence performance as compared to other activation functions</li>\n</ul>\n<p><strong>Disadvantage of ReLU</strong> - the dying ReLU problem. A ReLU neuron is dead if it is stuck in the negative side and always outputs zero (so it is impossible for the neuron to recover back)</p>\n<p><strong>To fix the ReLU disadvantage</strong> in 2017 Google Brain invented the <code>Swish</code> activation - almost <code>ReLU</code>, but with fixed negative part and the parameter to change the slope for the positives. It shows better scores on the different benchmarks like ImageNet and everybody was happy about that.</p>\n<p>But still researchers wanted <strong>better accuracy, stability and universality</strong> so in 2019 the <code>Mish</code> activation appears. Why it can be better than <code>Swish</code>?</p>\n<ul>\n<li><code>Mish</code> outperforms in case of noisy input conditions in comparison with <code>ReLU</code> and <code>Swish</code></li>\n<li>Test accuracy with different optimizers has less drop in case of <code>Mish</code> compared with <code>Swish</code></li>\n<li><code>Mish</code> has shown to have a consistent improvement over <code>Swish</code> using different dropout rates</li>\n<li>Test accuracy with different learning rates for <code>Mish</code> is also usually better than <code>Swish</code></li>\n</ul>\n<p>So that's why <code>Mish</code> become the better, more stable and universal activation for NNs but it is also important to say - <strong>to find the optimal activation function for the specific ML task/competition, you need to check different variants and compare them on your validation set</strong>. The data and the experiment should say what is the best choice, not the papers.</p>\n<p>Hope this helps you to make a better choice :)</p>\n<p>Alex</p>\n<p>P.S. For more details and graphs you can check <a href=\"https://krutikabapat.github.io/Swish-Vs-Mish-Latest-Activation-Functions/\" target=\"_blank\">this link</a></p>",
      "rawMarkdown": "Hi Kaggle fellows, \n\nIn this thread I would like to discuss the differences in the NN activation functions, as currently we have [the top-1 kernel](https://www.kaggle.com/code/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta) based on the NN built with FastAI framework.\n\nWhen you try to create the better NN model, you almost always try to find the best NN structure you can - and this is not only the number of layers and their sizes! It's also **important to choose the right activation for the layers** to prevent the usual NN problems like vanishing or exploding gradients, bad training speed or final quality.  \n\nWhat activations do we know, how do they look like in comparison to each other?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fac63fa4bb56c678383c0b68a163c5f4e%2Factivation.png?generation=1665783746972764&alt=media) \n\nThe most commonly known variants are Sigmoid, Tanh and ReLU, but all of them have some advantages and disadvantages. Let's take a closer look for the ReLU, the most common default activation for the frameworks nowadays.\n\n**Advantages of ReLU:**\n- For `x > 0` ReLU has constant gradient, which reduces the chance of vanishing gradients at any point of time (results in faster learning)\n- For `x <= 0` the gradient is zero and hence less number of neurons will be fired which reduce overfitting and is also cost efficient\n- ReLU has better convergence performance as compared to other activation functions\n\n**Disadvantage of ReLU** - the dying ReLU problem. A ReLU neuron is dead if it is stuck in the negative side and always outputs zero (so it is impossible for the neuron to recover back)\n\n**To fix the ReLU disadvantage** in 2017 Google Brain invented the `Swish` activation - almost `ReLU`, but with fixed negative part and the parameter to change the slope for the positives. It shows better scores on the different benchmarks like ImageNet and everybody was happy about that.\n\nBut still researchers wanted **better accuracy, stability and universality** so in 2019 the `Mish` activation appears. Why it can be better than `Swish`?\n- `Mish` outperforms in case of noisy input conditions in comparison with `ReLU` and `Swish`\n- Test accuracy with different optimizers has less drop in case of `Mish` compared with `Swish`\n- `Mish` has shown to have a consistent improvement over `Swish` using different dropout rates\n- Test accuracy with different learning rates for `Mish` is also usually better than `Swish`\n\nSo that's why `Mish` become the better, more stable and universal activation for NNs but it is also important to say - **to find the optimal activation function for the specific ML task/competition, you need to check different variants and compare them on your validation set**. The data and the experiment should say what is the best choice, not the papers.\n\nHope this helps you to make a better choice :)\n\nAlex\n\nP.S. For more details and graphs you can check [this link](https://krutikabapat.github.io/Swish-Vs-Mish-Latest-Activation-Functions/)",
      "votes": null
    },
    {
      "id": "1988514",
      "postDate": "10/15/2022 11:10:11",
      "content": "<p>Great post! In my experience, the choice of activation functions don't have a a huge impact for shallow neural networks (i.e. NNs with a small number of layers, e.g. less than 4) but it becomes very important for ultra deep neural networks (50+ layers) where you have the problem of exploding and vanishing gradients. I haven't heard of mish before though so thanks for sharing!</p>",
      "rawMarkdown": "Great post! In my experience, the choice of activation functions don't have a a huge impact for shallow neural networks (i.e. NNs with a small number of layers, e.g. less than 4) but it becomes very important for ultra deep neural networks (50+ layers) where you have the problem of exploding and vanishing gradients. I haven't heard of mish before though so thanks for sharing!",
      "votes": null
    },
    {
      "id": "1991099",
      "postDate": "10/17/2022 01:25:19",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "1991327",
      "postDate": "10/17/2022 05:28:20",
      "content": "<p>Good post! It helps me a lot.</p>",
      "rawMarkdown": "Good post! It helps me a lot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1988514,
      "author_name": "samuelcortinhas",
      "author_url": "",
      "post_date": "10/15/2022 11:10:11",
      "content": "<p>Great post! In my experience, the choice of activation functions don't have a a huge impact for shallow neural networks (i.e. NNs with a small number of layers, e.g. less than 4) but it becomes very important for ultra deep neural networks (50+ layers) where you have the problem of exploding and vanishing gradients. I haven't heard of mish before though so thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1991099,
      "author_name": "pomiro",
      "author_url": "",
      "post_date": "10/17/2022 01:25:19",
      "content": "<p>Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1991327,
      "author_name": "zhuoranjordan",
      "author_url": "",
      "post_date": "10/17/2022 05:28:20",
      "content": "<p>Good post! It helps me a lot.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1987904": "Hi Kaggle fellows, \n\nIn this thread I would like to discuss the differences in the NN activation functions, as currently we have [the top-1 kernel](https://www.kaggle.com/code/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta) based on the NN built with FastAI framework.\n\nWhen you try to create the better NN model, you almost always try to find the best NN structure you can - and this is not only the number of layers and their sizes! It's also **important to choose the right activation for the layers** to prevent the usual NN problems like vanishing or exploding gradients, bad training speed or final quality.  \n\nWhat activations do we know, how do they look like in comparison to each other?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fac63fa4bb56c678383c0b68a163c5f4e%2Factivation.png?generation=1665783746972764&alt=media) \n\nThe most commonly known variants are Sigmoid, Tanh and ReLU, but all of them have some advantages and disadvantages. Let's take a closer look for the ReLU, the most common default activation for the frameworks nowadays.\n\n**Advantages of ReLU:**\n- For `x > 0` ReLU has constant gradient, which reduces the chance of vanishing gradients at any point of time (results in faster learning)\n- For `x <= 0` the gradient is zero and hence less number of neurons will be fired which reduce overfitting and is also cost efficient\n- ReLU has better convergence performance as compared to other activation functions\n\n**Disadvantage of ReLU** - the dying ReLU problem. A ReLU neuron is dead if it is stuck in the negative side and always outputs zero (so it is impossible for the neuron to recover back)\n\n**To fix the ReLU disadvantage** in 2017 Google Brain invented the `Swish` activation - almost `ReLU`, but with fixed negative part and the parameter to change the slope for the positives. It shows better scores on the different benchmarks like ImageNet and everybody was happy about that.\n\nBut still researchers wanted **better accuracy, stability and universality** so in 2019 the `Mish` activation appears. Why it can be better than `Swish`?\n- `Mish` outperforms in case of noisy input conditions in comparison with `ReLU` and `Swish`\n- Test accuracy with different optimizers has less drop in case of `Mish` compared with `Swish`\n- `Mish` has shown to have a consistent improvement over `Swish` using different dropout rates\n- Test accuracy with different learning rates for `Mish` is also usually better than `Swish`\n\nSo that's why `Mish` become the better, more stable and universal activation for NNs but it is also important to say - **to find the optimal activation function for the specific ML task/competition, you need to check different variants and compare them on your validation set**. The data and the experiment should say what is the best choice, not the papers.\n\nHope this helps you to make a better choice :)\n\nAlex\n\nP.S. For more details and graphs you can check [this link](https://krutikabapat.github.io/Swish-Vs-Mish-Latest-Activation-Functions/)",
    "1988514": "Great post! In my experience, the choice of activation functions don't have a a huge impact for shallow neural networks (i.e. NNs with a small number of layers, e.g. less than 4) but it becomes very important for ultra deep neural networks (50+ layers) where you have the problem of exploding and vanishing gradients. I haven't heard of mish before though so thanks for sharing!",
    "1991099": "Thank you!",
    "1991327": "Good post! It helps me a lot."
  },
  "source": "meta"
}