{
  "id": 167764,
  "title": "Uncertainty prediction with neural nets",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/167764",
  "author_name": "kkiller",
  "post_date": "2020-07-17T18:00:46.305000",
  "votes": 34,
  "comment_count": 7,
  "views": 0,
  "content": "<p>This <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression\" target=\"_blank\">OSIC</a> competition is very special as it differs from traditional competitions which just ask for pointwise prediction. Here, we need to output a <strong>confidence level</strong> aside with our pointwise prediction. Those predictions are evaluated using <strong>laplace loglikelihood</strong> which is  a non classical metric.</p>\n<p>Those \"non classical\" things make this competition more interesting :). From the begining I went for neural nets (not exactly from the beginning as I've tried some bayesian regressions and I poorly failed to beat -7.xx). The main question remains how can we get a confidence level using a neural nets  ?</p>\n<h1>1. First attempt</h1>\n<blockquote>\n  <p>… let my architecture output 2 things (y<em>predict, y</em>confidence)</p>\n</blockquote>\n<p>This is the most obvious approach and that is where I started from. But as the loss is  a little bit weird, optimizing the neural nets using those outputs make my model quickly stucking in <strong>local minimums</strong>. The best score I got was around -8.xx . Perhaps I miss some regularisation over there but, instead of spending my time on stabilizing my model, I move from another approach …</p>\n<h1>2. Second attempt</h1>\n<blockquote>\n  <p>… building a smooth proxy for the loss function</p>\n</blockquote>\n<p>Laplace loglikelihood is explicit and could be written as sqrt<em>2xDelta/sigma + log(sqrt</em>2xsigma). Here by differentiating according to <strong>sigma</strong>, it comes that the loss would be optimal if <strong>sigma = sqrt_2xDelta</strong>. This observation was crucial when building my new loss which makes move from -8.xx to -7.0x</p>\n<h1>3. Third attempt</h1>\n<blockquote>\n  <p>… quantile regression and Pinball loss</p>\n</blockquote>\n<p>To be continued ….</p>",
  "messages": [
    {
      "id": 933463,
      "postDate": "2020-07-17T18:00:46.307Z",
      "content": "<p>This <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression\" target=\"_blank\">OSIC</a> competition is very special as it differs from traditional competitions which just ask for pointwise prediction. Here, we need to output a <strong>confidence level</strong> aside with our pointwise prediction. Those predictions are evaluated using <strong>laplace loglikelihood</strong> which is  a non classical metric.</p>\n<p>Those \"non classical\" things make this competition more interesting :). From the begining I went for neural nets (not exactly from the beginning as I've tried some bayesian regressions and I poorly failed to beat -7.xx). The main question remains how can we get a confidence level using a neural nets  ?</p>\n<h1>1. First attempt</h1>\n<blockquote>\n  <p>… let my architecture output 2 things (y<em>predict, y</em>confidence)</p>\n</blockquote>\n<p>This is the most obvious approach and that is where I started from. But as the loss is  a little bit weird, optimizing the neural nets using those outputs make my model quickly stucking in <strong>local minimums</strong>. The best score I got was around -8.xx . Perhaps I miss some regularisation over there but, instead of spending my time on stabilizing my model, I move from another approach …</p>\n<h1>2. Second attempt</h1>\n<blockquote>\n  <p>… building a smooth proxy for the loss function</p>\n</blockquote>\n<p>Laplace loglikelihood is explicit and could be written as sqrt<em>2xDelta/sigma + log(sqrt</em>2xsigma). Here by differentiating according to <strong>sigma</strong>, it comes that the loss would be optimal if <strong>sigma = sqrt_2xDelta</strong>. This observation was crucial when building my new loss which makes move from -8.xx to -7.0x</p>\n<h1>3. Third attempt</h1>\n<blockquote>\n  <p>… quantile regression and Pinball loss</p>\n</blockquote>\n<p>To be continued ….</p>",
      "rawMarkdown": "This [OSIC](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression) competition is very special as it differs from traditional competitions which just ask for pointwise prediction. Here, we need to output a **confidence level** aside with our pointwise prediction. Those predictions are evaluated using **laplace loglikelihood** which is  a non classical metric.\n\nThose \"non classical\" things make this competition more interesting :). From the begining I went for neural nets (not exactly from the beginning as I've tried some bayesian regressions and I poorly failed to beat -7.xx). The main question remains how can we get a confidence level using a neural nets  ?\n\n# 1. First attempt\n&gt;... let my architecture output 2 things (y_predict, y_confidence)\n\nThis is the most obvious approach and that is where I started from. But as the loss is  a little bit weird, optimizing the neural nets using those outputs make my model quickly stucking in **local minimums**. The best score I got was around -8.xx . Perhaps I miss some regularisation over there but, instead of spending my time on stabilizing my model, I move from another approach ...\n\n# 2. Second attempt\n&gt;... building a smooth proxy for the loss function\n\nLaplace loglikelihood is explicit and could be written as sqrt_2xDelta/sigma + log(sqrt_2xsigma). Here by differentiating according to **sigma**, it comes that the loss would be optimal if **sigma = sqrt_2xDelta**. This observation was crucial when building my new loss which makes move from -8.xx to -7.0x\n\n# 3. Third attempt\n&gt; ... quantile regression and Pinball loss\n\nTo be continued ....\n",
      "votes": 34
    },
    {
      "id": 933648,
      "postDate": "2020-07-17T21:21:38.830Z",
      "content": "<p>I need to learn about this pinball loss as I'm seeing it everywhere and it seems to be working well. </p>\n\n<p>On neural networks and uncertainty prediction I've written up a Bayesian neural network in pyro: <a href=\"https://www.kaggle.com/jameschapman19/bayesian-nn-pyro-tabular\">https://www.kaggle.com/jameschapman19/bayesian-nn-pyro-tabular</a></p>\n\n<p>Where I take confidence to be the standard deviation of the sampled outputs (I'm assuming your Bayes regression was similar). Someone mentioned on here that confidence in the Laplace likelihood is standard deviation.</p>\n\n<p>The pyro Bayesian neural network is something I've wanted to learn for a while and I'm glad this competition gave me the nudge!</p>",
      "rawMarkdown": "I need to learn about this pinball loss as I'm seeing it everywhere and it seems to be working well. \n\nOn neural networks and uncertainty prediction I've written up a Bayesian neural network in pyro: https://www.kaggle.com/jameschapman19/bayesian-nn-pyro-tabular\n\nWhere I take confidence to be the standard deviation of the sampled outputs (I'm assuming your Bayes regression was similar). Someone mentioned on here that confidence in the Laplace likelihood is standard deviation.\n\nThe pyro Bayesian neural network is something I've wanted to learn for a while and I'm glad this competition gave me the nudge!",
      "votes": 3,
      "replies": [
        {
          "id": 933710,
          "postDate": "2020-07-18T00:00:34.940Z",
          "content": "<p>Yes I will be updating the discussion by soon … Your notebook seems great, it proposes some different approach when compared to the main public kernels 👍</p>",
          "rawMarkdown": "Yes I will be updating the discussion by soon ... Your notebook seems great, it proposes some different approach when compared to the main public kernels 👍",
          "votes": 1
        },
        {
          "id": 937263,
          "postDate": "2020-07-20T21:42:11.947Z",
          "content": "<p>Had a stab at both the dropout bayesian approximation (dropout at test time) and a bootstrap ensemble of NNs (not sure this route will be the one for this competition given the relatively few subjects/observations). Dropout showed some promise actually given I didn't do any tuning. Guessing given your 8th ranking was the quantile regression stuff? <a href=\"https://www.kaggle.com/jameschapman19/dropout-as-bayesian-estimation\">https://www.kaggle.com/jameschapman19/dropout-as-bayesian-estimation</a></p>",
          "rawMarkdown": "Had a stab at both the dropout bayesian approximation (dropout at test time) and a bootstrap ensemble of NNs (not sure this route will be the one for this competition given the relatively few subjects/observations). Dropout showed some promise actually given I didn't do any tuning. Guessing given your 8th ranking was the quantile regression stuff? https://www.kaggle.com/jameschapman19/dropout-as-bayesian-estimation"
        }
      ]
    },
    {
      "id": 942439,
      "postDate": "2020-07-23T18:36:19.933Z",
      "content": "<p>Nice insight !!!</p>",
      "rawMarkdown": "Nice insight !!!",
      "votes": 2
    },
    {
      "id": 985549,
      "postDate": "2020-08-25T20:07:40.843Z",
      "content": "<p>in your second attempt, optimizing \\((\\sigma-\\sqrt(2)*\\Delta)^2\\) doesn't only bring \\(\\sigma\\) closer to \\(\\Delta\\) but also brings \\(\\Delta\\) closer to \\(\\sigma\\) because both of them are dependant on the network. So that would slow down the training.</p>",
      "rawMarkdown": "in your second attempt, optimizing \\\\((\\\\sigma-\\\\sqrt(2)*\\\\Delta)^2\\\\) doesn't only bring \\\\(\\\\sigma\\\\) closer to \\\\(\\\\Delta\\\\) but also brings \\\\(\\\\Delta\\\\) closer to \\\\(\\\\sigma\\\\) because both of them are dependant on the network. So that would slow down the training."
    },
    {
      "id": 981614,
      "postDate": "2020-08-22T15:33:16.093Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 981613,
      "postDate": "2020-08-22T15:33:16.087Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 933648,
      "author_name": "jameschapman19",
      "author_url": "",
      "post_date": "2020-07-17T21:21:38.830000",
      "content": "<p>I need to learn about this pinball loss as I'm seeing it everywhere and it seems to be working well. </p>\n\n<p>On neural networks and uncertainty prediction I've written up a Bayesian neural network in pyro: <a href=\"https://www.kaggle.com/jameschapman19/bayesian-nn-pyro-tabular\">https://www.kaggle.com/jameschapman19/bayesian-nn-pyro-tabular</a></p>\n\n<p>Where I take confidence to be the standard deviation of the sampled outputs (I'm assuming your Bayes regression was similar). Someone mentioned on here that confidence in the Laplace likelihood is standard deviation.</p>\n\n<p>The pyro Bayesian neural network is something I've wanted to learn for a while and I'm glad this competition gave me the nudge!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 933710,
          "author_name": "kkiller",
          "author_url": "",
          "post_date": "2020-07-18T00:00:34.940000",
          "content": "<p>Yes I will be updating the discussion by soon … Your notebook seems great, it proposes some different approach when compared to the main public kernels 👍</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 937263,
          "author_name": "jameschapman19",
          "author_url": "",
          "post_date": "2020-07-20T21:42:11.947000",
          "content": "<p>Had a stab at both the dropout bayesian approximation (dropout at test time) and a bootstrap ensemble of NNs (not sure this route will be the one for this competition given the relatively few subjects/observations). Dropout showed some promise actually given I didn't do any tuning. Guessing given your 8th ranking was the quantile regression stuff? <a href=\"https://www.kaggle.com/jameschapman19/dropout-as-bayesian-estimation\">https://www.kaggle.com/jameschapman19/dropout-as-bayesian-estimation</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 942439,
      "author_name": "Ulrich G.",
      "author_url": "",
      "post_date": "2020-07-23T18:36:19.933000",
      "content": "<p>Nice insight !!!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 985549,
      "author_name": "DarkCube",
      "author_url": "",
      "post_date": "2020-08-25T20:07:40.843000",
      "content": "<p>in your second attempt, optimizing \\((\\sigma-\\sqrt(2)*\\Delta)^2\\) doesn't only bring \\(\\sigma\\) closer to \\(\\Delta\\) but also brings \\(\\Delta\\) closer to \\(\\sigma\\) because both of them are dependant on the network. So that would slow down the training.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 981614,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-22T15:33:16.093000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 981613,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-22T15:33:16.087000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "933463": "This [OSIC](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression) competition is very special as it differs from traditional competitions which just ask for pointwise prediction. Here, we need to output a **confidence level** aside with our pointwise prediction. Those predictions are evaluated using **laplace loglikelihood** which is  a non classical metric.\n\nThose \"non classical\" things make this competition more interesting :). From the begining I went for neural nets (not exactly from the beginning as I've tried some bayesian regressions and I poorly failed to beat -7.xx). The main question remains how can we get a confidence level using a neural nets  ?\n\n# 1. First attempt\n&gt;... let my architecture output 2 things (y_predict, y_confidence)\n\nThis is the most obvious approach and that is where I started from. But as the loss is  a little bit weird, optimizing the neural nets using those outputs make my model quickly stucking in **local minimums**. The best score I got was around -8.xx . Perhaps I miss some regularisation over there but, instead of spending my time on stabilizing my model, I move from another approach ...\n\n# 2. Second attempt\n&gt;... building a smooth proxy for the loss function\n\nLaplace loglikelihood is explicit and could be written as sqrt_2xDelta/sigma + log(sqrt_2xsigma). Here by differentiating according to **sigma**, it comes that the loss would be optimal if **sigma = sqrt_2xDelta**. This observation was crucial when building my new loss which makes move from -8.xx to -7.0x\n\n# 3. Third attempt\n&gt; ... quantile regression and Pinball loss\n\nTo be continued ....\n",
    "933648": "I need to learn about this pinball loss as I'm seeing it everywhere and it seems to be working well. \n\nOn neural networks and uncertainty prediction I've written up a Bayesian neural network in pyro: https://www.kaggle.com/jameschapman19/bayesian-nn-pyro-tabular\n\nWhere I take confidence to be the standard deviation of the sampled outputs (I'm assuming your Bayes regression was similar). Someone mentioned on here that confidence in the Laplace likelihood is standard deviation.\n\nThe pyro Bayesian neural network is something I've wanted to learn for a while and I'm glad this competition gave me the nudge!",
    "942439": "Nice insight !!!",
    "985549": "in your second attempt, optimizing \\\\((\\\\sigma-\\\\sqrt(2)*\\\\Delta)^2\\\\) doesn't only bring \\\\(\\\\sigma\\\\) closer to \\\\(\\\\Delta\\\\) but also brings \\\\(\\\\Delta\\\\) closer to \\\\(\\\\sigma\\\\) because both of them are dependant on the network. So that would slow down the training.",
    "981614": "",
    "981613": ""
  }
}