{
  "id": 39206,
  "title": "Different Activation",
  "url": "/competitions/carvana-image-masking-challenge/discussion/39206",
  "author_name": "",
  "post_date": "2017-09-09T08:22:59.334631900Z",
  "votes": 2,
  "comment_count": 13,
  "views": 0,
  "content": "<p>As seen in these two pictures, the activation<code>PReLu</code>  preforms better than <code>ReLu</code>.\nHowever, my memory is limited when I try the <code>PReLu</code> activation. Has anyone tried different activation?</p>",
  "messages": [
    {
      "id": "219706",
      "postDate": "09/09/2017 08:22:59",
      "content": "<p>As seen in these two pictures, the activation<code>PReLu</code>  preforms better than <code>ReLu</code>.\nHowever, my memory is limited when I try the <code>PReLu</code> activation. Has anyone tried different activation?</p>",
      "rawMarkdown": "As seen in these two pictures, the activation```PReLu```  preforms better than ```ReLu```.\nHowever, my memory is limited when I try the ```PReLu``` activation. Has anyone tried different activation?",
      "votes": null
    },
    {
      "id": "219778",
      "postDate": "09/09/2017 16:16:06",
      "content": "<p>PReLu <em>might</em> perform better than ReLu, but there is additional memory overhead. <em>Most</em> of the time it isn't a big deal, but if you're already using most of your VRAM it can be tight.</p>",
      "rawMarkdown": "PReLu *might* perform better than ReLu, but there is additional memory overhead. *Most* of the time it isn't a big deal, but if you're already using most of your VRAM it can be tight.",
      "votes": null
    },
    {
      "id": "219904",
      "postDate": "09/10/2017 08:24:45",
      "content": "<p><code>relu</code>:</p>\n\n<p>dice_coeff: 0.9877  val_dice_coeff: 0.9918</p>\n\n<p><code>PReLU</code>:</p>\n\n<p>dice_coeff: 0.9877  val_dice_coeff: 0.9919</p>",
      "rawMarkdown": "```relu```:\n\ndice_coeff: 0.9877  val_dice_coeff: 0.9918\n\n ```PReLU```:\n\ndice_coeff: 0.9877  val_dice_coeff: 0.9919",
      "votes": null
    },
    {
      "id": "219914",
      "postDate": "09/10/2017 09:04:00",
      "content": "<p>if prelu works, you may also want to try elu. My guess there will be differences, but not very big. Since all the results are about the same, you can ensemble them together later.</p>",
      "rawMarkdown": "if prelu works, you may also want to try elu. My guess there will be differences, but not very big. Since all the results are about the same, you can ensemble them together later.",
      "votes": null
    },
    {
      "id": "219916",
      "postDate": "09/10/2017 09:14:24",
      "content": "<p>Yes, I am trying elu.\nThank you for your advice.</p>",
      "rawMarkdown": "Yes, I am trying elu.\nThank you for your advice.",
      "votes": null
    },
    {
      "id": "219973",
      "postDate": "09/10/2017 15:13:06",
      "content": "<p>some experiments with input 128*128</p>\n\n<p>elu:\nval_loss : 0.128</p>\n\n<p>relu:\nval_loss : 0.129</p>\n\n<p>selu:\nval_loss : 0.131</p>",
      "rawMarkdown": "some experiments with input 128*128\n\nelu:\nval_loss : 0.128\n\nrelu:\nval_loss : 0.129\n\nselu:\nval_loss : 0.131",
      "votes": null
    },
    {
      "id": "220063",
      "postDate": "09/11/2017 03:12:08",
      "content": "<p>@atom1231 is that selu and removing bn layers?</p>",
      "rawMarkdown": "atom1231 is that selu and removing bn layers?",
      "votes": null
    },
    {
      "id": "220087",
      "postDate": "09/11/2017 05:32:29",
      "content": "<p>yes</p>",
      "rawMarkdown": "yes",
      "votes": null
    },
    {
      "id": "220139",
      "postDate": "09/11/2017 10:19:42",
      "content": "<p>Did you make sure mean=0, var=1 and used 'lecun_normal' as init? Both are required to make SELU work.</p>",
      "rawMarkdown": "Did you make sure mean=0, var=1 and used 'lecun_normal' as init? Both are required to make SELU work.",
      "votes": null
    },
    {
      "id": "220150",
      "postDate": "09/11/2017 10:54:38",
      "content": "<p>I am using 'he_normal' as init.</p>",
      "rawMarkdown": "I am using 'he_normal' as init.",
      "votes": null
    },
    {
      "id": "220164",
      "postDate": "09/11/2017 11:35:43",
      "content": "<p>In the SELU paper:</p>\n\n<p>Initialization. Since SNNs have a fixed point at zero mean and unit variance for normalized weights\nω = Pn i=1 wi = 0 and τ = Pn i=1 w2 i = 1 (see above), we initialize SNNs such that these constraints\nare fulfilled in expectation. We draw the weights from a Gaussian distribution with E(wi) = 0 and\nvariance Var(wi) = 1/n. Uniform and truncated Gaussian distributions with these moments led to\nnetworks with similar behavior. The “MSRA initialization” is similar since it uses zero mean and\nvariance 2/n to initialize the weights [17]. The additional factor 2 counters the effect of rectified\nlinear units.</p>\n\n<p>Can you try lecun_normal?</p>",
      "rawMarkdown": "In the SELU paper:\n\nInitialization. Since SNNs have a fixed point at zero mean and unit variance for normalized weights\nω = Pn i=1 wi = 0 and τ = Pn i=1 w2 i = 1 (see above), we initialize SNNs such that these constraints\nare fulfilled in expectation. We draw the weights from a Gaussian distribution with E(wi) = 0 and\nvariance Var(wi) = 1/n. Uniform and truncated Gaussian distributions with these moments led to\nnetworks with similar behavior. The “MSRA initialization” is similar since it uses zero mean and\nvariance 2/n to initialize the weights [17]. The additional factor 2 counters the effect of rectified\nlinear units.\n\nCan you try lecun_normal?",
      "votes": null
    },
    {
      "id": "220165",
      "postDate": "09/11/2017 11:38:24",
      "content": "<p>Sorry , I haven't tried SELU. I only tried relu, prelu and elu. And the elu activation performed better. </p>",
      "rawMarkdown": "Sorry , I haven't tried SELU. I only tried relu, prelu and elu. And the elu activation performed better.",
      "votes": null
    },
    {
      "id": "220176",
      "postDate": "09/11/2017 12:18:43",
      "content": "<p>@Andres <br>\nActually i did not follow all the selu assumption to do the experiment . just do a simple try based on current architecture .\nSince I did not  find any positive report/example from internet  about selu with cnn/unet , I stop to explore the topic.</p>",
      "rawMarkdown": "Andres  \nActually i did not follow all the selu assumption to do the experiment . just do a simple try based on current architecture .\nSince I did not  find any positive report/example from internet  about selu with cnn/unet , I stop to explore the topic.",
      "votes": null
    },
    {
      "id": "220749",
      "postDate": "09/13/2017 00:39:19",
      "content": "<p>Yes, so I have tried elu, but the LB score didn't make any change comparing to relu. So I will try to ensemble them later.</p>",
      "rawMarkdown": "Yes, so I have tried elu, but the LB score didn't make any change comparing to relu. So I will try to ensemble them later.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 219778,
      "author_name": "stevenknguyen",
      "author_url": "",
      "post_date": "09/09/2017 16:16:06",
      "content": "<p>PReLu <em>might</em> perform better than ReLu, but there is additional memory overhead. <em>Most</em> of the time it isn't a big deal, but if you're already using most of your VRAM it can be tight.</p>",
      "votes": null,
      "replies": [
        {
          "id": 220749,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "09/13/2017 00:39:19",
          "content": "<p>Yes, so I have tried elu, but the LB score didn't make any change comparing to relu. So I will try to ensemble them later.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 219904,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "09/10/2017 08:24:45",
      "content": "<p><code>relu</code>:</p>\n\n<p>dice_coeff: 0.9877  val_dice_coeff: 0.9918</p>\n\n<p><code>PReLU</code>:</p>\n\n<p>dice_coeff: 0.9877  val_dice_coeff: 0.9919</p>",
      "votes": null,
      "replies": [
        {
          "id": 219914,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/10/2017 09:04:00",
          "content": "<p>if prelu works, you may also want to try elu. My guess there will be differences, but not very big. Since all the results are about the same, you can ensemble them together later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 219916,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "09/10/2017 09:14:24",
          "content": "<p>Yes, I am trying elu.\nThank you for your advice.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 219973,
          "author_name": "atom1231",
          "author_url": "",
          "post_date": "09/10/2017 15:13:06",
          "content": "<p>some experiments with input 128*128</p>\n\n<p>elu:\nval_loss : 0.128</p>\n\n<p>relu:\nval_loss : 0.129</p>\n\n<p>selu:\nval_loss : 0.131</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 220063,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "09/11/2017 03:12:08",
          "content": "<p>@atom1231 is that selu and removing bn layers?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 220087,
          "author_name": "atom1231",
          "author_url": "",
          "post_date": "09/11/2017 05:32:29",
          "content": "<p>yes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 220139,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "09/11/2017 10:19:42",
          "content": "<p>Did you make sure mean=0, var=1 and used 'lecun_normal' as init? Both are required to make SELU work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 220150,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "09/11/2017 10:54:38",
          "content": "<p>I am using 'he_normal' as init.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 220164,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "09/11/2017 11:35:43",
          "content": "<p>In the SELU paper:</p>\n\n<p>Initialization. Since SNNs have a fixed point at zero mean and unit variance for normalized weights\nω = Pn i=1 wi = 0 and τ = Pn i=1 w2 i = 1 (see above), we initialize SNNs such that these constraints\nare fulfilled in expectation. We draw the weights from a Gaussian distribution with E(wi) = 0 and\nvariance Var(wi) = 1/n. Uniform and truncated Gaussian distributions with these moments led to\nnetworks with similar behavior. The “MSRA initialization” is similar since it uses zero mean and\nvariance 2/n to initialize the weights [17]. The additional factor 2 counters the effect of rectified\nlinear units.</p>\n\n<p>Can you try lecun_normal?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 220165,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "09/11/2017 11:38:24",
          "content": "<p>Sorry , I haven't tried SELU. I only tried relu, prelu and elu. And the elu activation performed better. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 220176,
          "author_name": "atom1231",
          "author_url": "",
          "post_date": "09/11/2017 12:18:43",
          "content": "<p>@Andres <br>\nActually i did not follow all the selu assumption to do the experiment . just do a simple try based on current architecture .\nSince I did not  find any positive report/example from internet  about selu with cnn/unet , I stop to explore the topic.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "219706": "As seen in these two pictures, the activation```PReLu```  preforms better than ```ReLu```.\nHowever, my memory is limited when I try the ```PReLu``` activation. Has anyone tried different activation?",
    "219778": "PReLu *might* perform better than ReLu, but there is additional memory overhead. *Most* of the time it isn't a big deal, but if you're already using most of your VRAM it can be tight.",
    "219904": "```relu```:\n\ndice_coeff: 0.9877  val_dice_coeff: 0.9918\n\n ```PReLU```:\n\ndice_coeff: 0.9877  val_dice_coeff: 0.9919",
    "219914": "if prelu works, you may also want to try elu. My guess there will be differences, but not very big. Since all the results are about the same, you can ensemble them together later.",
    "219916": "Yes, I am trying elu.\nThank you for your advice.",
    "219973": "some experiments with input 128*128\n\nelu:\nval_loss : 0.128\n\nrelu:\nval_loss : 0.129\n\nselu:\nval_loss : 0.131",
    "220063": "atom1231 is that selu and removing bn layers?",
    "220087": "yes",
    "220139": "Did you make sure mean=0, var=1 and used 'lecun_normal' as init? Both are required to make SELU work.",
    "220150": "I am using 'he_normal' as init.",
    "220164": "In the SELU paper:\n\nInitialization. Since SNNs have a fixed point at zero mean and unit variance for normalized weights\nω = Pn i=1 wi = 0 and τ = Pn i=1 w2 i = 1 (see above), we initialize SNNs such that these constraints\nare fulfilled in expectation. We draw the weights from a Gaussian distribution with E(wi) = 0 and\nvariance Var(wi) = 1/n. Uniform and truncated Gaussian distributions with these moments led to\nnetworks with similar behavior. The “MSRA initialization” is similar since it uses zero mean and\nvariance 2/n to initialize the weights [17]. The additional factor 2 counters the effect of rectified\nlinear units.\n\nCan you try lecun_normal?",
    "220165": "Sorry , I haven't tried SELU. I only tried relu, prelu and elu. And the elu activation performed better.",
    "220176": "Andres  \nActually i did not follow all the selu assumption to do the experiment . just do a simple try based on current architecture .\nSince I did not  find any positive report/example from internet  about selu with cnn/unet , I stop to explore the topic.",
    "220749": "Yes, so I have tried elu, but the LB score didn't make any change comparing to relu. So I will try to ensemble them later."
  },
  "source": "meta"
}