{
  "id": 82696,
  "title": "NN regression: ReLU vs clipping?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/82696",
  "author_name": "",
  "post_date": "2019-03-03T17:45:16.599488300Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Since time to failure (ttf) is always &gt; 0, what is the appropriate way to enforce this behaviour with neural networks? I've tried using ReLU activation on the output neuron or using no activation and clipping all values smaller than zero in post-processing. \nI think relu might lead to overall higher ttf values and probably more problems for low ttf, but seems like the correct way of implementation (besides similar activations). \nCan you only find the better option by trial and error or is there any guideline to such problems? Since i couldn't make any conclusion with my relatively simple model, has anyone investigated this?</p>",
  "messages": [
    {
      "id": "482822",
      "postDate": "03/03/2019 17:45:16",
      "content": "<p>Since time to failure (ttf) is always &gt; 0, what is the appropriate way to enforce this behaviour with neural networks? I've tried using ReLU activation on the output neuron or using no activation and clipping all values smaller than zero in post-processing. \nI think relu might lead to overall higher ttf values and probably more problems for low ttf, but seems like the correct way of implementation (besides similar activations). \nCan you only find the better option by trial and error or is there any guideline to such problems? Since i couldn't make any conclusion with my relatively simple model, has anyone investigated this?</p>",
      "rawMarkdown": "Since time to failure (ttf) is always &gt; 0, what is the appropriate way to enforce this behaviour with neural networks? I've tried using ReLU activation on the output neuron or using no activation and clipping all values smaller than zero in post-processing. \nI think relu might lead to overall higher ttf values and probably more problems for low ttf, but seems like the correct way of implementation (besides similar activations). \nCan you only find the better option by trial and error or is there any guideline to such problems? Since i couldn't make any conclusion with my relatively simple model, has anyone investigated this?",
      "votes": null
    },
    {
      "id": "483248",
      "postDate": "03/04/2019 11:21:44",
      "content": "<p>I personally would not add any activation at the end of the network. I don't know if there's a big chance that the network will output negative values because it doesn't see any during the training. But even if, you can clip it during the prediction. </p>\n\n<p>I think using ReLU during the training can slow down the learning on early stages when the weights can be distracted not knowing that they should output a positive value per se. Nevertheless, the difference should not be big.</p>",
      "rawMarkdown": "I personally would not add any activation at the end of the network. I don't know if there's a big chance that the network will output negative values because it doesn't see any during the training. But even if, you can clip it during the prediction. \n\nI think using ReLU during the training can slow down the learning on early stages when the weights can be distracted not knowing that they should output a positive value per se. Nevertheless, the difference should not be big.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 483248,
      "author_name": "davids1992",
      "author_url": "",
      "post_date": "03/04/2019 11:21:44",
      "content": "<p>I personally would not add any activation at the end of the network. I don't know if there's a big chance that the network will output negative values because it doesn't see any during the training. But even if, you can clip it during the prediction. </p>\n\n<p>I think using ReLU during the training can slow down the learning on early stages when the weights can be distracted not knowing that they should output a positive value per se. Nevertheless, the difference should not be big.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "482822": "Since time to failure (ttf) is always &gt; 0, what is the appropriate way to enforce this behaviour with neural networks? I've tried using ReLU activation on the output neuron or using no activation and clipping all values smaller than zero in post-processing. \nI think relu might lead to overall higher ttf values and probably more problems for low ttf, but seems like the correct way of implementation (besides similar activations). \nCan you only find the better option by trial and error or is there any guideline to such problems? Since i couldn't make any conclusion with my relatively simple model, has anyone investigated this?",
    "483248": "I personally would not add any activation at the end of the network. I don't know if there's a big chance that the network will output negative values because it doesn't see any during the training. But even if, you can clip it during the prediction. \n\nI think using ReLU during the training can slow down the learning on early stages when the weights can be distracted not knowing that they should output a positive value per se. Nevertheless, the difference should not be big."
  },
  "source": "meta"
}