{
  "id": 18352,
  "title": "Clinical Significance",
  "url": "/competitions/second-annual-data-science-bowl/discussion/18352",
  "author_name": "",
  "post_date": "2016-01-11T21:40:48.923Z",
  "votes": 4,
  "comment_count": 10,
  "views": 1926,
  "content": "<p>Based on the recent <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18346/answers-q-a-with-principle-investigators-michael-hansen-ph-d-and-dr-andrew\">Q&amp;A</a> with Dr Hansen and Dr Arai,  it appears that in order to be clinically useful a model should be accurate to within 20 ml and preferably within 10 ml. The question is then: how does this relate to the CRPS scores used to score this contest?</p>\n\n<p>If one takes a naive approach and models these predictions as a step function, then a prediction that was off by 10 ml would have a CRPS of 0.0167. However, one can do considerably better than just predicting step functions as has been discussed in this <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17966/optimal-submission-for-continuous-ranked-probability-score-crps\">thread</a>. Based on some slightly dubious back of the envelope calculations, I'm estimating that the error for the optimal distribution is in the vicinity of \\(0.5 \\sigma / 600\\) where \\(\\sigma\\) is the RMS error of the prediction.  <em>This estimate is dodgy and I encourage anyone to improve it; either through math or simulation. Note also that the relationship between the RMS error and CRPS of the optimal distribution is probably dependent on the error distribution of the prediction.</em></p>\n\n<p>Based on the above estimate we get a CRPS score of 0.008 for an RMS error of 10 ml and 0.016 for an RMS error of 20 ml. The latter definitely seems achievable, but the former is going to be a challenge. That gives us a goal to shoot for however: 0.008 or bust!</p>",
  "messages": [
    {
      "id": "104336",
      "postDate": "01/11/2016 21:40:48",
      "content": "<p>Based on the recent <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18346/answers-q-a-with-principle-investigators-michael-hansen-ph-d-and-dr-andrew\">Q&amp;A</a> with Dr Hansen and Dr Arai,  it appears that in order to be clinically useful a model should be accurate to within 20 ml and preferably within 10 ml. The question is then: how does this relate to the CRPS scores used to score this contest?</p>\n\n<p>If one takes a naive approach and models these predictions as a step function, then a prediction that was off by 10 ml would have a CRPS of 0.0167. However, one can do considerably better than just predicting step functions as has been discussed in this <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17966/optimal-submission-for-continuous-ranked-probability-score-crps\">thread</a>. Based on some slightly dubious back of the envelope calculations, I'm estimating that the error for the optimal distribution is in the vicinity of \\(0.5 \\sigma / 600\\) where \\(\\sigma\\) is the RMS error of the prediction.  <em>This estimate is dodgy and I encourage anyone to improve it; either through math or simulation. Note also that the relationship between the RMS error and CRPS of the optimal distribution is probably dependent on the error distribution of the prediction.</em></p>\n\n<p>Based on the above estimate we get a CRPS score of 0.008 for an RMS error of 10 ml and 0.016 for an RMS error of 20 ml. The latter definitely seems achievable, but the former is going to be a challenge. That gives us a goal to shoot for however: 0.008 or bust!</p>",
      "rawMarkdown": "Based on the recent [Q&A][1] with Dr Hansen and Dr Arai,  it appears that in order to be clinically useful a model should be accurate to within 20 ml and preferably within 10 ml. The question is then: how does this relate to the CRPS scores used to score this contest?\r\n\r\nIf one takes a naive approach and models these predictions as a step function, then a prediction that was off by 10 ml would have a CRPS of 0.0167. However, one can do considerably better than just predicting step functions as has been discussed in this [thread][2]. Based on some slightly dubious back of the envelope calculations, I'm estimating that the error for the optimal distribution is in the vicinity of \\\\(0.5 \\sigma / 600\\\\) where \\\\(\\sigma\\\\) is the RMS error of the prediction.  *This estimate is dodgy and I encourage anyone to improve it; either through math or simulation. Note also that the relationship between the RMS error and CRPS of the optimal distribution is probably dependent on the error distribution of the prediction.*\r\n\r\nBased on the above estimate we get a CRPS score of 0.008 for an RMS error of 10 ml and 0.016 for an RMS error of 20 ml. The latter definitely seems achievable, but the former is going to be a challenge. That gives us a goal to shoot for however: 0.008 or bust!\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18346/answers-q-a-with-principle-investigators-michael-hansen-ph-d-and-dr-andrew\r\n  [2]: https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17966/optimal-submission-for-continuous-ranked-probability-score-crps",
      "votes": null
    },
    {
      "id": "104346",
      "postDate": "01/11/2016 23:45:10",
      "content": "<p>Here's another, probably more sensible, approach to this:  let's assume that our predictions follow a Gaussian distribution with a standard deviation \\(\\sigma\\).  The CDF associated with this is given by:\n$$\n\\text{CDF}(v) = \\frac{1}{2}\\text{erfc}\\left(\\frac{v_0 - v}{\\sqrt{2}\\sigma} \\right)\n$$\nwhere \\(v_0\\) is the true volume and erfc is the complementary error function (see <a href=\"http://www.wolframalpha.com/input/?i=normal distribution&lk=4&num=1\">here</a>).</p>\n\n<p>We can compute the CRPS of this as:\n$$\n\\text{CRPS} = \\int_{-\\infty}^{\\infty}\\left(\\text{h}(v-v_0) - \\text{CDF}(v) \\right)^2 =  2\\int_{-\\infty}^{v_0}  {\\text{CDF}(v)}^2\n$$\ndue to the symmetry of the problem. This <a href=\"http://www.wolframalpha.com/input/?i=+2+*+int_0^inf+(1+-+erfc(-x%2Fsqrt(2))+%2F+2)^2\">evaluates</a> to \\(0.234\\sigma\\).</p>\n\n<p>That's over a factor of two lower than my previous estimate, which is unfortunate. While a score of 0.008 seems within reach, a score of 0.004 seems likely to be very challenging!  I'm hoping someone can find something wrong with this (for example, a nice factor of two which would make this agree with the earlier estimate). </p>\n\n<p>[EDIT] It turns out that the links above are broken. Here they are on their own:\n <a href=\"http://www.wolframalpha.com/input/?i=normal distribution&lk=4&num=1\">http://www.wolframalpha.com/input/?i=normal%20distribution&amp;lk=4&amp;num=1</a>\n <a href=\"http://www.wolframalpha.com/input/?i=+2+\">http://www.wolframalpha.com/input/?i=+2+</a>*+int_0%5Einf+%281+-+erfc%28-x%2Fsqrt%282%29%29+%2F+2%29%5E2</p>\n\n<p>[EDIT2] Added missing zeros (0.04-&gt;0.004, 0.08-&gt;0.008)</p>",
      "rawMarkdown": "Here's another, probably more sensible, approach to this:  let's assume that our predictions follow a Gaussian distribution with a standard deviation \\\\(\\sigma\\\\).  The CDF associated with this is given by:\r\n$$\r\n\\text{CDF}(v) = \\frac{1}{2}\\text{erfc}\\left(\\frac{v_0 - v}{\\sqrt{2}\\sigma} \\right)\r\n$$\r\nwhere \\\\(v_0\\\\) is the true volume and erfc is the complementary error function (see [here][1]).\r\n\r\nWe can compute the CRPS of this as:\r\n$$\r\n\\text{CRPS} = \\int_{-\\infty}^{\\infty}\\left(\\text{h}(v-v_0) - \\text{CDF}(v) \\right)^2 =  2\\int_{-\\infty}^{v_0}  {\\text{CDF}(v)}^2\r\n$$\r\ndue to the symmetry of the problem. This [evaluates][2] to \\\\(0.234\\sigma\\\\).\r\n\r\nThat's over a factor of two lower than my previous estimate, which is unfortunate. While a score of 0.008 seems within reach, a score of 0.004 seems likely to be very challenging!  I'm hoping someone can find something wrong with this (for example, a nice factor of two which would make this agree with the earlier estimate). \r\n\r\n[EDIT] It turns out that the links above are broken. Here they are on their own:\r\n http://www.wolframalpha.com/input/?i=normal%20distribution&lk=4&num=1\r\n http://www.wolframalpha.com/input/?i=+2+*+int_0%5Einf+%281+-+erfc%28-x%2Fsqrt%282%29%29+%2F+2%29%5E2\r\n\r\n[EDIT2] Added missing zeros (0.04->0.004, 0.08->0.008)\r\n\r\n  [1]: http://www.wolframalpha.com/input/?i=normal%20distribution&lk=4&num=1\r\n  [2]: http://www.wolframalpha.com/input/?i=+2+*+int_0%5Einf+%281+-+erfc%28-x%2Fsqrt%282%29%29+%2F+2%29%5E2",
      "votes": null
    },
    {
      "id": "104359",
      "postDate": "01/12/2016 04:08:05",
      "content": "<p>I spent some time manually drawing some contours by hand and calculated the final volumes to check my backend math and get a feel for where it gets difficult to predict the technician's choice of boundary.  Once I got the hang of it I was matching +-8mL from the training value.</p>\n\n<p>So I think the repeatability can be there.  And I would say below +-10mL the difference in contours is small enough that it's hard to say who's contour is correct.</p>",
      "rawMarkdown": "I spent some time manually drawing some contours by hand and calculated the final volumes to check my backend math and get a feel for where it gets difficult to predict the technician's choice of boundary.  Once I got the hang of it I was matching +-8mL from the training value.\r\n\r\nSo I think the repeatability can be there.  And I would say below +-10mL the difference in contours is small enough that it's hard to say who's contour is correct.",
      "votes": null
    },
    {
      "id": "104367",
      "postDate": "01/12/2016 07:00:45",
      "content": "<p>@Tim</p>\n\n<p>Thanks for commenting on the scoring issue. I think that the choice of CRPS for scoring instead of a simple sums of absolute error values is somewhat counterproductive for the purpose of obtaining the most accurate measurements. The reason is that inferior quality solution can evaluate better than the superior one, having all errors smaller, only due to selection of cleverer probability distribution functions (see <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17926/problems-with-evaluation-description/104287#post104287\">my post here</a>).</p>\n\n<p>BTW, you meant 0.008 and 0.004 two posts up, right? </p>\n\n<blockquote>\n  <p><em>While a score of 0.08 seems within reach, a score of 0.04 seems likely to be very challenging</em></p>\n</blockquote>",
      "rawMarkdown": "Tim\r\n\r\nThanks for commenting on the scoring issue. I think that the choice of CRPS for scoring instead of a simple sums of absolute error values is somewhat counterproductive for the purpose of obtaining the most accurate measurements. The reason is that inferior quality solution can evaluate better than the superior one, having all errors smaller, only due to selection of cleverer probability distribution functions (see [my post here][1]).\r\n\r\nBTW, you meant 0.008 and 0.004 two posts up, right? \r\n\r\n> *While a score of 0.08 seems within reach, a score of 0.04 seems likely to be very challenging*\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17926/problems-with-evaluation-description/104287#post104287",
      "votes": null
    },
    {
      "id": "104381",
      "postDate": "01/12/2016 10:17:34",
      "content": "<p>Hi Tim,</p>\n\n<p>Thank you for your comment and examples provided above!</p>\n\n<p>From what i think, I guess the \\(v_0\\) should be the mean of distribution of a particular prediction, instead of the true volume.</p>\n\n<p>$$\n\\text{CDF}(v) = \\frac{1}{2}\\text{erfc}\\left(\\frac{v_0 - v}{\\sqrt{2}\\sigma} \\right)\n$$</p>\n\n<p>I think if we let \\(v_1\\) to be the true volume, CRPS will be:</p>\n\n<p>$$\n\\text{CRPS} = \\int_{-\\infty}^{\\infty}\\left(\\text{h}(v-v_1) - \\text{CDF}(v) \\right)^2\n$$\n$$\n= \\int_{-\\infty}^{v_1}\\left(\\text{CDF}(v) \\right)^2 + \\int_{v_1}^{\\infty}\\left(\\text{1} -  \\text{CDF}(v) \\right)^2\n$$</p>\n\n<p>due to symmetry, here assume \\(v_0\\), mean of the predicted distribution is smaller than the actual volume \\(v_1\\).</p>\n\n<p>I tried to work out the calculation, assuming \\(v_1\\) is one standard deviation away from \\(v_0\\),</p>\n\n<p><a href=\"http://www.wolframalpha.com/input/?i=integrate[(0.5+\">http://www.wolframalpha.com/input/?i=integrate%5B%280.5+</a>*+erfc%28-x%2Fsqrt%282%29%29+%29%5E2%2C+%7Bx%2C-inf%2C0%7D%5D</p>\n\n<p><a href=\"http://www.wolframalpha.com/input/?i=integrate[(0.5+\">http://www.wolframalpha.com/input/?i=integrate%5B%280.5+</a>*+erfc%28-x%2Fsqrt%282%29%29+%29%5E2%2C+%7Bx%2C0%2C1%7D%5D</p>\n\n<p><a href=\"http://www.wolframalpha.com/input/?i=integrate[(1+-+0.5+\">http://www.wolframalpha.com/input/?i=integrate%5B%281+-+0.5+</a>*+erfc%28-x%2Fsqrt%282%29%29+%29%5E2%2C+%7Bx%2C1%2C1000%7D%5D</p>\n\n<p>(sorry for the multiple links, I couldn't figure out a better way to work around the wolfram integration XD)</p>\n\n<p>The CRPS for this instance is (0.6024\\sigma). Let me know if I got it wrong in these. </p>\n\n<p>(i had another try with \\(v_1\\) two standard deviation away from \\(v_0\\), the CRPS is (1.45\\sigma)</p>",
      "rawMarkdown": "Hi Tim,\r\n\r\nThank you for your comment and examples provided above!\r\n\r\nFrom what i think, I guess the \\\\(v_0\\\\) should be the mean of distribution of a particular prediction, instead of the true volume.\r\n\r\n$$\r\n\\text{CDF}(v) = \\frac{1}{2}\\text{erfc}\\left(\\frac{v_0 - v}{\\sqrt{2}\\sigma} \\right)\r\n$$\r\n\r\nI think if we let \\\\(v_1\\\\) to be the true volume, CRPS will be:\r\n\r\n$$\r\n\\text{CRPS} = \\int_{-\\infty}^{\\infty}\\left(\\text{h}(v-v_1) - \\text{CDF}(v) \\right)^2\r\n$$\r\n$$\r\n= \\int_{-\\infty}^{v_1}\\left(\\text{CDF}(v) \\right)^2 + \\int_{v_1}^{\\infty}\\left(\\text{1} -  \\text{CDF}(v) \\right)^2\r\n$$\r\n\r\ndue to symmetry, here assume \\\\(v_0\\\\), mean of the predicted distribution is smaller than the actual volume \\\\(v_1\\\\).\r\n\r\nI tried to work out the calculation, assuming \\\\(v_1\\\\) is one standard deviation away from \\\\(v_0\\\\),\r\n\r\nhttp://www.wolframalpha.com/input/?i=integrate%5B%280.5+*+erfc%28-x%2Fsqrt%282%29%29+%29%5E2%2C+%7Bx%2C-inf%2C0%7D%5D\r\n\r\nhttp://www.wolframalpha.com/input/?i=integrate%5B%280.5+*+erfc%28-x%2Fsqrt%282%29%29+%29%5E2%2C+%7Bx%2C0%2C1%7D%5D\r\n\r\nhttp://www.wolframalpha.com/input/?i=integrate%5B%281+-+0.5+*+erfc%28-x%2Fsqrt%282%29%29+%29%5E2%2C+%7Bx%2C1%2C1000%7D%5D\r\n\r\n(sorry for the multiple links, I couldn't figure out a better way to work around the wolfram integration XD)\r\n\r\nThe CRPS for this instance is \\(0.6024\\sigma\\). Let me know if I got it wrong in these. \r\n\r\n(i had another try with \\\\(v_1\\\\) two standard deviation away from \\\\(v_0\\\\), the CRPS is \\(1.45\\sigma\\)",
      "votes": null
    },
    {
      "id": "104398",
      "postDate": "01/12/2016 13:11:52",
      "content": "<p>@athyssen , that's interesting that you can easily hit +-10 ml by hand. That would suggest that training a model to segment the individual slices and adding them up would be successful. However, my understanding is -- and I admit I didn't look into this approach very much -- is that there is a very limited amount of ground truth data for segmenting the slices individually, which makes that approach difficult to make competitive.  </p>",
      "rawMarkdown": "athyssen , that's interesting that you can easily hit +-10 ml by hand. That would suggest that training a model to segment the individual slices and adding them up would be successful. However, my understanding is -- and I admit I didn't look into this approach very much -- is that there is a very limited amount of ground truth data for segmenting the slices individually, which makes that approach difficult to make competitive.",
      "votes": null
    },
    {
      "id": "104403",
      "postDate": "01/12/2016 13:38:20",
      "content": "<p>@Paul Jurczak, thanks for the correction. </p>\n\n<p>CRPS seems to have both a plus and a minus. On the plus side, it incorporates some notion of confidence into the model. The more confident the model is in a given prediction, the narrower the CDF it can spit out. So, for instance if some data is noisy, or otherwise hard to predict from, it can spit out a wide CDF.  On the negative side is interpretability: how does this measure relate back to +-10-20 ml that the doctors say would be useful? That second problem is what I'm trying to address here.</p>\n\n<p>As for the issue of a more accurate model being beaten by a poorer model due to the choice of CDF, that doesn't worry me to much.  If you have a model that predicts volumes fairly accurately, it shouldn't be too difficult to turn it into a &quot;good&quot; CDF. I suspect that just using <strong>cerf</strong> as used above should work pretty well. Just tune \\(\\sigma\\) to get the best score you can on your local validation data.  Another approach would be to just steal the back end from the MXNet example, but instead of feeding it images, just feed it your predictions. That would turn volume predictions into optimal or near-optimal CDFs.</p>",
      "rawMarkdown": "Paul Jurczak, thanks for the correction. \r\n\r\nCRPS seems to have both a plus and a minus. On the plus side, it incorporates some notion of confidence into the model. The more confident the model is in a given prediction, the narrower the CDF it can spit out. So, for instance if some data is noisy, or otherwise hard to predict from, it can spit out a wide CDF.  On the negative side is interpretability: how does this measure relate back to +-10-20 ml that the doctors say would be useful? That second problem is what I'm trying to address here.\r\n\r\nAs for the issue of a more accurate model being beaten by a poorer model due to the choice of CDF, that doesn't worry me to much.  If you have a model that predicts volumes fairly accurately, it shouldn't be too difficult to turn it into a \"good\" CDF. I suspect that just using **cerf** as used above should work pretty well. Just tune \\\\(\\sigma\\\\) to get the best score you can on your local validation data.  Another approach would be to just steal the back end from the MXNet example, but instead of feeding it images, just feed it your predictions. That would turn volume predictions into optimal or near-optimal CDFs.",
      "votes": null
    },
    {
      "id": "104404",
      "postDate": "01/12/2016 13:40:20",
      "content": "<p>@sakimilo, thanks for taking a look at the math. I'll look through what you've done and compare it to my stuff a little later -- after the coffee has had time to kick in!</p>",
      "rawMarkdown": "sakimilo, thanks for taking a look at the math. I'll look through what you've done and compare it to my stuff a little later -- after the coffee has had time to kick in!",
      "votes": null
    },
    {
      "id": "104429",
      "postDate": "01/12/2016 16:50:33",
      "content": "<p>@sakimilo,</p>\n\n<p>[Thinking &quot;out loud&quot;, so forgive me if I ramble a bit]</p>\n\n<p>My main interest here is connecting the prediction accuracy that is clinically relevant (\\(\\pm10\\text{-}20 \\text{ ml} \\)) to the leaderboard CRPS scores.   So when a doctor says that a predictor (either a model or a cardiologist) is accurate to within 10 ml, what does that mean? I'm going to assume that it means that the RMS error is 10 ml. I'm further going to assume,  to make things tractable, that the error is normally distributed with mean 0. If the mean is not zero, that means that the predictor has some sort of systematic errors,  and I'm not going to worry about systematic errors, at least for right now.  </p>\n\n<p>That's a long winded way of justifying the assumption that  \\(v_1 = v_0\\) above. Once that assumption is made, one can then go on to assume that \\(v_1 = v_0 = 0\\) since it won't affect the results of the integral.  At that point it's fairly easy to do the integral (and by that I mean have Wolfram Alpha do the integral for me).  </p>\n\n<p>I believe what you are doing above, and please correct me if I'm wrong,  is assuming specific systematic errors and then computing the CRPS for those values. That's interesting, but I think that it's difficult to get far thinking about systematic errors since they could have any structure.</p>",
      "rawMarkdown": "sakimilo,\r\n\r\n[Thinking \"out loud\", so forgive me if I ramble a bit]\r\n\r\nMy main interest here is connecting the prediction accuracy that is clinically relevant (\\\\(\\pm10\\text{-}20 \\text{ ml} \\\\)) to the leaderboard CRPS scores.   So when a doctor says that a predictor (either a model or a cardiologist) is accurate to within 10 ml, what does that mean? I'm going to assume that it means that the RMS error is 10 ml. I'm further going to assume,  to make things tractable, that the error is normally distributed with mean 0. If the mean is not zero, that means that the predictor has some sort of systematic errors,  and I'm not going to worry about systematic errors, at least for right now.  \r\n\r\nThat's a long winded way of justifying the assumption that  \\\\(v_1 = v_0\\\\) above. Once that assumption is made, one can then go on to assume that \\\\(v_1 = v_0 = 0\\\\) since it won't affect the results of the integral.  At that point it's fairly easy to do the integral (and by that I mean have Wolfram Alpha do the integral for me).  \r\n\r\nI believe what you are doing above, and please correct me if I'm wrong,  is assuming specific systematic errors and then computing the CRPS for those values. That's interesting, but I think that it's difficult to get far thinking about systematic errors since they could have any structure.",
      "votes": null
    },
    {
      "id": "104753",
      "postDate": "01/16/2016 09:35:04",
      "content": "<p>@Tim Hochberg</p>\n\n<p>Can you comment on why you relate CRPS to RMSE, but not MAE (Mean Absolute Error)? I did some variational calculus which suggest me it is a better indicator. </p>",
      "rawMarkdown": "Tim Hochberg\r\n\r\nCan you comment on why you relate CRPS to RMSE, but not MAE (Mean Absolute Error)? I did some variational calculus which suggest me it is a better indicator.",
      "votes": null
    },
    {
      "id": "104923",
      "postDate": "01/18/2016 01:07:59",
      "content": "<p>@bobye,  mainly because RMSE is easier to work with. I'm not trying to make any sort of precise calculation, I'm just trying to get a general idea of what CRPS score corresponds to, for example, &quot;+-10 ml&quot;. Now, without further specification &quot;+-10 ml&quot; could mean many things, so I chose to treat it as RMSE since that's typically the easiest to work with. </p>",
      "rawMarkdown": "bobye,  mainly because RMSE is easier to work with. I'm not trying to make any sort of precise calculation, I'm just trying to get a general idea of what CRPS score corresponds to, for example, \"+-10 ml\". Now, without further specification \"+-10 ml\" could mean many things, so I chose to treat it as RMSE since that's typically the easiest to work with.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 104346,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "01/11/2016 23:45:10",
      "content": "<p>Here's another, probably more sensible, approach to this:  let's assume that our predictions follow a Gaussian distribution with a standard deviation \\(\\sigma\\).  The CDF associated with this is given by:\n$$\n\\text{CDF}(v) = \\frac{1}{2}\\text{erfc}\\left(\\frac{v_0 - v}{\\sqrt{2}\\sigma} \\right)\n$$\nwhere \\(v_0\\) is the true volume and erfc is the complementary error function (see <a href=\"http://www.wolframalpha.com/input/?i=normal distribution&lk=4&num=1\">here</a>).</p>\n\n<p>We can compute the CRPS of this as:\n$$\n\\text{CRPS} = \\int_{-\\infty}^{\\infty}\\left(\\text{h}(v-v_0) - \\text{CDF}(v) \\right)^2 =  2\\int_{-\\infty}^{v_0}  {\\text{CDF}(v)}^2\n$$\ndue to the symmetry of the problem. This <a href=\"http://www.wolframalpha.com/input/?i=+2+*+int_0^inf+(1+-+erfc(-x%2Fsqrt(2))+%2F+2)^2\">evaluates</a> to \\(0.234\\sigma\\).</p>\n\n<p>That's over a factor of two lower than my previous estimate, which is unfortunate. While a score of 0.008 seems within reach, a score of 0.004 seems likely to be very challenging!  I'm hoping someone can find something wrong with this (for example, a nice factor of two which would make this agree with the earlier estimate). </p>\n\n<p>[EDIT] It turns out that the links above are broken. Here they are on their own:\n <a href=\"http://www.wolframalpha.com/input/?i=normal distribution&lk=4&num=1\">http://www.wolframalpha.com/input/?i=normal%20distribution&amp;lk=4&amp;num=1</a>\n <a href=\"http://www.wolframalpha.com/input/?i=+2+\">http://www.wolframalpha.com/input/?i=+2+</a>*+int_0%5Einf+%281+-+erfc%28-x%2Fsqrt%282%29%29+%2F+2%29%5E2</p>\n\n<p>[EDIT2] Added missing zeros (0.04-&gt;0.004, 0.08-&gt;0.008)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104359,
      "author_name": "athyssen",
      "author_url": "",
      "post_date": "01/12/2016 04:08:05",
      "content": "<p>I spent some time manually drawing some contours by hand and calculated the final volumes to check my backend math and get a feel for where it gets difficult to predict the technician's choice of boundary.  Once I got the hang of it I was matching +-8mL from the training value.</p>\n\n<p>So I think the repeatability can be there.  And I would say below +-10mL the difference in contours is small enough that it's hard to say who's contour is correct.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104367,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "01/12/2016 07:00:45",
      "content": "<p>@Tim</p>\n\n<p>Thanks for commenting on the scoring issue. I think that the choice of CRPS for scoring instead of a simple sums of absolute error values is somewhat counterproductive for the purpose of obtaining the most accurate measurements. The reason is that inferior quality solution can evaluate better than the superior one, having all errors smaller, only due to selection of cleverer probability distribution functions (see <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17926/problems-with-evaluation-description/104287#post104287\">my post here</a>).</p>\n\n<p>BTW, you meant 0.008 and 0.004 two posts up, right? </p>\n\n<blockquote>\n  <p><em>While a score of 0.08 seems within reach, a score of 0.04 seems likely to be very challenging</em></p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104381,
      "author_name": "sakimilo",
      "author_url": "",
      "post_date": "01/12/2016 10:17:34",
      "content": "<p>Hi Tim,</p>\n\n<p>Thank you for your comment and examples provided above!</p>\n\n<p>From what i think, I guess the \\(v_0\\) should be the mean of distribution of a particular prediction, instead of the true volume.</p>\n\n<p>$$\n\\text{CDF}(v) = \\frac{1}{2}\\text{erfc}\\left(\\frac{v_0 - v}{\\sqrt{2}\\sigma} \\right)\n$$</p>\n\n<p>I think if we let \\(v_1\\) to be the true volume, CRPS will be:</p>\n\n<p>$$\n\\text{CRPS} = \\int_{-\\infty}^{\\infty}\\left(\\text{h}(v-v_1) - \\text{CDF}(v) \\right)^2\n$$\n$$\n= \\int_{-\\infty}^{v_1}\\left(\\text{CDF}(v) \\right)^2 + \\int_{v_1}^{\\infty}\\left(\\text{1} -  \\text{CDF}(v) \\right)^2\n$$</p>\n\n<p>due to symmetry, here assume \\(v_0\\), mean of the predicted distribution is smaller than the actual volume \\(v_1\\).</p>\n\n<p>I tried to work out the calculation, assuming \\(v_1\\) is one standard deviation away from \\(v_0\\),</p>\n\n<p><a href=\"http://www.wolframalpha.com/input/?i=integrate[(0.5+\">http://www.wolframalpha.com/input/?i=integrate%5B%280.5+</a>*+erfc%28-x%2Fsqrt%282%29%29+%29%5E2%2C+%7Bx%2C-inf%2C0%7D%5D</p>\n\n<p><a href=\"http://www.wolframalpha.com/input/?i=integrate[(0.5+\">http://www.wolframalpha.com/input/?i=integrate%5B%280.5+</a>*+erfc%28-x%2Fsqrt%282%29%29+%29%5E2%2C+%7Bx%2C0%2C1%7D%5D</p>\n\n<p><a href=\"http://www.wolframalpha.com/input/?i=integrate[(1+-+0.5+\">http://www.wolframalpha.com/input/?i=integrate%5B%281+-+0.5+</a>*+erfc%28-x%2Fsqrt%282%29%29+%29%5E2%2C+%7Bx%2C1%2C1000%7D%5D</p>\n\n<p>(sorry for the multiple links, I couldn't figure out a better way to work around the wolfram integration XD)</p>\n\n<p>The CRPS for this instance is (0.6024\\sigma). Let me know if I got it wrong in these. </p>\n\n<p>(i had another try with \\(v_1\\) two standard deviation away from \\(v_0\\), the CRPS is (1.45\\sigma)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104398,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "01/12/2016 13:11:52",
      "content": "<p>@athyssen , that's interesting that you can easily hit +-10 ml by hand. That would suggest that training a model to segment the individual slices and adding them up would be successful. However, my understanding is -- and I admit I didn't look into this approach very much -- is that there is a very limited amount of ground truth data for segmenting the slices individually, which makes that approach difficult to make competitive.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104403,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "01/12/2016 13:38:20",
      "content": "<p>@Paul Jurczak, thanks for the correction. </p>\n\n<p>CRPS seems to have both a plus and a minus. On the plus side, it incorporates some notion of confidence into the model. The more confident the model is in a given prediction, the narrower the CDF it can spit out. So, for instance if some data is noisy, or otherwise hard to predict from, it can spit out a wide CDF.  On the negative side is interpretability: how does this measure relate back to +-10-20 ml that the doctors say would be useful? That second problem is what I'm trying to address here.</p>\n\n<p>As for the issue of a more accurate model being beaten by a poorer model due to the choice of CDF, that doesn't worry me to much.  If you have a model that predicts volumes fairly accurately, it shouldn't be too difficult to turn it into a &quot;good&quot; CDF. I suspect that just using <strong>cerf</strong> as used above should work pretty well. Just tune \\(\\sigma\\) to get the best score you can on your local validation data.  Another approach would be to just steal the back end from the MXNet example, but instead of feeding it images, just feed it your predictions. That would turn volume predictions into optimal or near-optimal CDFs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104404,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "01/12/2016 13:40:20",
      "content": "<p>@sakimilo, thanks for taking a look at the math. I'll look through what you've done and compare it to my stuff a little later -- after the coffee has had time to kick in!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104429,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "01/12/2016 16:50:33",
      "content": "<p>@sakimilo,</p>\n\n<p>[Thinking &quot;out loud&quot;, so forgive me if I ramble a bit]</p>\n\n<p>My main interest here is connecting the prediction accuracy that is clinically relevant (\\(\\pm10\\text{-}20 \\text{ ml} \\)) to the leaderboard CRPS scores.   So when a doctor says that a predictor (either a model or a cardiologist) is accurate to within 10 ml, what does that mean? I'm going to assume that it means that the RMS error is 10 ml. I'm further going to assume,  to make things tractable, that the error is normally distributed with mean 0. If the mean is not zero, that means that the predictor has some sort of systematic errors,  and I'm not going to worry about systematic errors, at least for right now.  </p>\n\n<p>That's a long winded way of justifying the assumption that  \\(v_1 = v_0\\) above. Once that assumption is made, one can then go on to assume that \\(v_1 = v_0 = 0\\) since it won't affect the results of the integral.  At that point it's fairly easy to do the integral (and by that I mean have Wolfram Alpha do the integral for me).  </p>\n\n<p>I believe what you are doing above, and please correct me if I'm wrong,  is assuming specific systematic errors and then computing the CRPS for those values. That's interesting, but I think that it's difficult to get far thinking about systematic errors since they could have any structure.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104753,
      "author_name": "amazingbob",
      "author_url": "",
      "post_date": "01/16/2016 09:35:04",
      "content": "<p>@Tim Hochberg</p>\n\n<p>Can you comment on why you relate CRPS to RMSE, but not MAE (Mean Absolute Error)? I did some variational calculus which suggest me it is a better indicator. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104923,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "01/18/2016 01:07:59",
      "content": "<p>@bobye,  mainly because RMSE is easier to work with. I'm not trying to make any sort of precise calculation, I'm just trying to get a general idea of what CRPS score corresponds to, for example, &quot;+-10 ml&quot;. Now, without further specification &quot;+-10 ml&quot; could mean many things, so I chose to treat it as RMSE since that's typically the easiest to work with. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "104336": "Based on the recent [Q&A][1] with Dr Hansen and Dr Arai,  it appears that in order to be clinically useful a model should be accurate to within 20 ml and preferably within 10 ml. The question is then: how does this relate to the CRPS scores used to score this contest?\r\n\r\nIf one takes a naive approach and models these predictions as a step function, then a prediction that was off by 10 ml would have a CRPS of 0.0167. However, one can do considerably better than just predicting step functions as has been discussed in this [thread][2]. Based on some slightly dubious back of the envelope calculations, I'm estimating that the error for the optimal distribution is in the vicinity of \\\\(0.5 \\sigma / 600\\\\) where \\\\(\\sigma\\\\) is the RMS error of the prediction.  *This estimate is dodgy and I encourage anyone to improve it; either through math or simulation. Note also that the relationship between the RMS error and CRPS of the optimal distribution is probably dependent on the error distribution of the prediction.*\r\n\r\nBased on the above estimate we get a CRPS score of 0.008 for an RMS error of 10 ml and 0.016 for an RMS error of 20 ml. The latter definitely seems achievable, but the former is going to be a challenge. That gives us a goal to shoot for however: 0.008 or bust!\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18346/answers-q-a-with-principle-investigators-michael-hansen-ph-d-and-dr-andrew\r\n  [2]: https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17966/optimal-submission-for-continuous-ranked-probability-score-crps",
    "104346": "Here's another, probably more sensible, approach to this:  let's assume that our predictions follow a Gaussian distribution with a standard deviation \\\\(\\sigma\\\\).  The CDF associated with this is given by:\r\n$$\r\n\\text{CDF}(v) = \\frac{1}{2}\\text{erfc}\\left(\\frac{v_0 - v}{\\sqrt{2}\\sigma} \\right)\r\n$$\r\nwhere \\\\(v_0\\\\) is the true volume and erfc is the complementary error function (see [here][1]).\r\n\r\nWe can compute the CRPS of this as:\r\n$$\r\n\\text{CRPS} = \\int_{-\\infty}^{\\infty}\\left(\\text{h}(v-v_0) - \\text{CDF}(v) \\right)^2 =  2\\int_{-\\infty}^{v_0}  {\\text{CDF}(v)}^2\r\n$$\r\ndue to the symmetry of the problem. This [evaluates][2] to \\\\(0.234\\sigma\\\\).\r\n\r\nThat's over a factor of two lower than my previous estimate, which is unfortunate. While a score of 0.008 seems within reach, a score of 0.004 seems likely to be very challenging!  I'm hoping someone can find something wrong with this (for example, a nice factor of two which would make this agree with the earlier estimate). \r\n\r\n[EDIT] It turns out that the links above are broken. Here they are on their own:\r\n http://www.wolframalpha.com/input/?i=normal%20distribution&lk=4&num=1\r\n http://www.wolframalpha.com/input/?i=+2+*+int_0%5Einf+%281+-+erfc%28-x%2Fsqrt%282%29%29+%2F+2%29%5E2\r\n\r\n[EDIT2] Added missing zeros (0.04->0.004, 0.08->0.008)\r\n\r\n  [1]: http://www.wolframalpha.com/input/?i=normal%20distribution&lk=4&num=1\r\n  [2]: http://www.wolframalpha.com/input/?i=+2+*+int_0%5Einf+%281+-+erfc%28-x%2Fsqrt%282%29%29+%2F+2%29%5E2",
    "104359": "I spent some time manually drawing some contours by hand and calculated the final volumes to check my backend math and get a feel for where it gets difficult to predict the technician's choice of boundary.  Once I got the hang of it I was matching +-8mL from the training value.\r\n\r\nSo I think the repeatability can be there.  And I would say below +-10mL the difference in contours is small enough that it's hard to say who's contour is correct.",
    "104367": "Tim\r\n\r\nThanks for commenting on the scoring issue. I think that the choice of CRPS for scoring instead of a simple sums of absolute error values is somewhat counterproductive for the purpose of obtaining the most accurate measurements. The reason is that inferior quality solution can evaluate better than the superior one, having all errors smaller, only due to selection of cleverer probability distribution functions (see [my post here][1]).\r\n\r\nBTW, you meant 0.008 and 0.004 two posts up, right? \r\n\r\n> *While a score of 0.08 seems within reach, a score of 0.04 seems likely to be very challenging*\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17926/problems-with-evaluation-description/104287#post104287",
    "104381": "Hi Tim,\r\n\r\nThank you for your comment and examples provided above!\r\n\r\nFrom what i think, I guess the \\\\(v_0\\\\) should be the mean of distribution of a particular prediction, instead of the true volume.\r\n\r\n$$\r\n\\text{CDF}(v) = \\frac{1}{2}\\text{erfc}\\left(\\frac{v_0 - v}{\\sqrt{2}\\sigma} \\right)\r\n$$\r\n\r\nI think if we let \\\\(v_1\\\\) to be the true volume, CRPS will be:\r\n\r\n$$\r\n\\text{CRPS} = \\int_{-\\infty}^{\\infty}\\left(\\text{h}(v-v_1) - \\text{CDF}(v) \\right)^2\r\n$$\r\n$$\r\n= \\int_{-\\infty}^{v_1}\\left(\\text{CDF}(v) \\right)^2 + \\int_{v_1}^{\\infty}\\left(\\text{1} -  \\text{CDF}(v) \\right)^2\r\n$$\r\n\r\ndue to symmetry, here assume \\\\(v_0\\\\), mean of the predicted distribution is smaller than the actual volume \\\\(v_1\\\\).\r\n\r\nI tried to work out the calculation, assuming \\\\(v_1\\\\) is one standard deviation away from \\\\(v_0\\\\),\r\n\r\nhttp://www.wolframalpha.com/input/?i=integrate%5B%280.5+*+erfc%28-x%2Fsqrt%282%29%29+%29%5E2%2C+%7Bx%2C-inf%2C0%7D%5D\r\n\r\nhttp://www.wolframalpha.com/input/?i=integrate%5B%280.5+*+erfc%28-x%2Fsqrt%282%29%29+%29%5E2%2C+%7Bx%2C0%2C1%7D%5D\r\n\r\nhttp://www.wolframalpha.com/input/?i=integrate%5B%281+-+0.5+*+erfc%28-x%2Fsqrt%282%29%29+%29%5E2%2C+%7Bx%2C1%2C1000%7D%5D\r\n\r\n(sorry for the multiple links, I couldn't figure out a better way to work around the wolfram integration XD)\r\n\r\nThe CRPS for this instance is \\(0.6024\\sigma\\). Let me know if I got it wrong in these. \r\n\r\n(i had another try with \\\\(v_1\\\\) two standard deviation away from \\\\(v_0\\\\), the CRPS is \\(1.45\\sigma\\)",
    "104398": "athyssen , that's interesting that you can easily hit +-10 ml by hand. That would suggest that training a model to segment the individual slices and adding them up would be successful. However, my understanding is -- and I admit I didn't look into this approach very much -- is that there is a very limited amount of ground truth data for segmenting the slices individually, which makes that approach difficult to make competitive.",
    "104403": "Paul Jurczak, thanks for the correction. \r\n\r\nCRPS seems to have both a plus and a minus. On the plus side, it incorporates some notion of confidence into the model. The more confident the model is in a given prediction, the narrower the CDF it can spit out. So, for instance if some data is noisy, or otherwise hard to predict from, it can spit out a wide CDF.  On the negative side is interpretability: how does this measure relate back to +-10-20 ml that the doctors say would be useful? That second problem is what I'm trying to address here.\r\n\r\nAs for the issue of a more accurate model being beaten by a poorer model due to the choice of CDF, that doesn't worry me to much.  If you have a model that predicts volumes fairly accurately, it shouldn't be too difficult to turn it into a \"good\" CDF. I suspect that just using **cerf** as used above should work pretty well. Just tune \\\\(\\sigma\\\\) to get the best score you can on your local validation data.  Another approach would be to just steal the back end from the MXNet example, but instead of feeding it images, just feed it your predictions. That would turn volume predictions into optimal or near-optimal CDFs.",
    "104404": "sakimilo, thanks for taking a look at the math. I'll look through what you've done and compare it to my stuff a little later -- after the coffee has had time to kick in!",
    "104429": "sakimilo,\r\n\r\n[Thinking \"out loud\", so forgive me if I ramble a bit]\r\n\r\nMy main interest here is connecting the prediction accuracy that is clinically relevant (\\\\(\\pm10\\text{-}20 \\text{ ml} \\\\)) to the leaderboard CRPS scores.   So when a doctor says that a predictor (either a model or a cardiologist) is accurate to within 10 ml, what does that mean? I'm going to assume that it means that the RMS error is 10 ml. I'm further going to assume,  to make things tractable, that the error is normally distributed with mean 0. If the mean is not zero, that means that the predictor has some sort of systematic errors,  and I'm not going to worry about systematic errors, at least for right now.  \r\n\r\nThat's a long winded way of justifying the assumption that  \\\\(v_1 = v_0\\\\) above. Once that assumption is made, one can then go on to assume that \\\\(v_1 = v_0 = 0\\\\) since it won't affect the results of the integral.  At that point it's fairly easy to do the integral (and by that I mean have Wolfram Alpha do the integral for me).  \r\n\r\nI believe what you are doing above, and please correct me if I'm wrong,  is assuming specific systematic errors and then computing the CRPS for those values. That's interesting, but I think that it's difficult to get far thinking about systematic errors since they could have any structure.",
    "104753": "Tim Hochberg\r\n\r\nCan you comment on why you relate CRPS to RMSE, but not MAE (Mean Absolute Error)? I did some variational calculus which suggest me it is a better indicator.",
    "104923": "bobye,  mainly because RMSE is easier to work with. I'm not trying to make any sort of precise calculation, I'm just trying to get a general idea of what CRPS score corresponds to, for example, \"+-10 ml\". Now, without further specification \"+-10 ml\" could mean many things, so I chose to treat it as RMSE since that's typically the easiest to work with."
  },
  "source": "meta"
}