{
  "id": 17989,
  "title": "On low scores and long decimals",
  "url": "/competitions/second-annual-data-science-bowl/discussion/17989",
  "author_name": "",
  "post_date": "2015-12-17T16:28:50.277Z",
  "votes": 10,
  "comment_count": 8,
  "views": 4737,
  "content": "<p>We've received some inquiries/concerns about the &quot;low&quot; early competition scores, enough to warrant a thread about what it means for the CRPS to be &quot;low&quot;.</p>\n\n<p>With any metric (but especially mathematically complex ones) it's important not to let intuition overwhelm your interpretation. It's easy to <em>feel</em> that 1 is high and 0 is low, and therefore the current scores are in danger of saturation. However, you should keep in mind three important points here:</p>\n\n<ol>\n<li>The exponent in the metric is squaring a number &lt; 1 (yielding an even smaller number).</li>\n<li>You can and should plot your probability curves on top of a vertical line for the volume. If you do this, you can visually see that there is a lot of room for improvement, even for CRPS scores that feel low.</li>\n<li>Medical use cases impose a very high bar on the level of robustness and accuracy for an algorithm to make it into the clinic.</li>\n</ol>\n\n<p>What's the takeaway? When you look at the leaderboard and see a number that feels low, or gaps that seem very close, view the plots and run the numbers before you get discouraged. It's a loooong road to get to 0.0!</p>",
  "messages": [
    {
      "id": "101878",
      "postDate": "12/17/2015 16:28:50",
      "content": "<p>We've received some inquiries/concerns about the &quot;low&quot; early competition scores, enough to warrant a thread about what it means for the CRPS to be &quot;low&quot;.</p>\n\n<p>With any metric (but especially mathematically complex ones) it's important not to let intuition overwhelm your interpretation. It's easy to <em>feel</em> that 1 is high and 0 is low, and therefore the current scores are in danger of saturation. However, you should keep in mind three important points here:</p>\n\n<ol>\n<li>The exponent in the metric is squaring a number &lt; 1 (yielding an even smaller number).</li>\n<li>You can and should plot your probability curves on top of a vertical line for the volume. If you do this, you can visually see that there is a lot of room for improvement, even for CRPS scores that feel low.</li>\n<li>Medical use cases impose a very high bar on the level of robustness and accuracy for an algorithm to make it into the clinic.</li>\n</ol>\n\n<p>What's the takeaway? When you look at the leaderboard and see a number that feels low, or gaps that seem very close, view the plots and run the numbers before you get discouraged. It's a loooong road to get to 0.0!</p>",
      "rawMarkdown": "We've received some inquiries/concerns about the \"low\" early competition scores, enough to warrant a thread about what it means for the CRPS to be \"low\".\r\n\r\nWith any metric (but especially mathematically complex ones) it's important not to let intuition overwhelm your interpretation. It's easy to *feel* that 1 is high and 0 is low, and therefore the current scores are in danger of saturation. However, you should keep in mind three important points here:\r\n\r\n 1. The exponent in the metric is squaring a number < 1 (yielding an even smaller number).\r\n 2. You can and should plot your probability curves on top of a vertical line for the volume. If you do this, you can visually see that there is a lot of room for improvement, even for CRPS scores that feel low.\r\n 3. Medical use cases impose a very high bar on the level of robustness and accuracy for an algorithm to make it into the clinic.\r\n\r\nWhat's the takeaway? When you look at the leaderboard and see a number that feels low, or gaps that seem very close, view the plots and run the numbers before you get discouraged. It's a loooong road to get to 0.0!",
      "votes": null
    },
    {
      "id": "104503",
      "postDate": "01/13/2016 11:59:08",
      "content": "<p>Thanks to Jason Farbman at Booz Allen for putting together this graphic. (<a href=\"http://www.datasciencebowl.com/crps_and_its_implications/\">Re-posted from the Data Science Bowl Blog.</a>)</p>\n\n<p><img src=\"http://www.datasciencebowl.com/wp-content/uploads/2016/01/booz-dsb-err-3.png\" alt=\"CRPS and its implications\" title></p>",
      "rawMarkdown": "Thanks to Jason Farbman at Booz Allen for putting together this graphic. ([Re-posted from the Data Science Bowl Blog.][2])\r\n\r\n\r\n ![CRPS and its implications][1]\r\n\r\n\r\n\r\n  [1]: http://www.datasciencebowl.com/wp-content/uploads/2016/01/booz-dsb-err-3.png\r\n  [2]: http://www.datasciencebowl.com/crps_and_its_implications/",
      "votes": null
    },
    {
      "id": "104646",
      "postDate": "01/14/2016 21:58:46",
      "content": "<p>CRPS is not very intuitive at first, but the more I explore it, the more I like it.</p>\n\n<p>Imho the key observation is that, once you made your best guesses, you'll score higher with a cumulative distribution that reflects the true variance instead of a step function. This forces competitors to reveal not only their best guesses, but also the amount of noise they believe would remain unexplained. </p>\n\n<p>Congrats admins for this very clever choice... fair and robust!</p>",
      "rawMarkdown": "CRPS is not very intuitive at first, but the more I explore it, the more I like it.\r\n\r\nImho the key observation is that, once you made your best guesses, you'll score higher with a cumulative distribution that reflects the true variance instead of a step function. This forces competitors to reveal not only their best guesses, but also the amount of noise they believe would remain unexplained. \r\n\r\nCongrats admins for this very clever choice... fair and robust!",
      "votes": null
    },
    {
      "id": "104655",
      "postDate": "01/14/2016 23:49:31",
      "content": "<p>@Shannon</p>\n\n<p>You've shown in your post how small volume measurement error can dramatically affect EF value. The problem I have with this example is that EF can be calculated much more accurately without relying on EDV and ESV. I think that you can obtain high precision measurement of LV contraction based only on clearly imaged slices and disregard the top and bottom ones, which are the most difficult to interpret and most likely contribute most of the volume measurement error.</p>\n\n<p>I wrote about it earlier, but I think that having EF as a part of submission score independent of EDV and ESV would lead to more clinically useful solution.</p>",
      "rawMarkdown": "Shannon\r\n\r\nYou've shown in your post how small volume measurement error can dramatically affect EF value. The problem I have with this example is that EF can be calculated much more accurately without relying on EDV and ESV. I think that you can obtain high precision measurement of LV contraction based only on clearly imaged slices and disregard the top and bottom ones, which are the most difficult to interpret and most likely contribute most of the volume measurement error.\r\n\r\nI wrote about it earlier, but I think that having EF as a part of submission score independent of EDV and ESV would lead to more clinically useful solution.",
      "votes": null
    },
    {
      "id": "104659",
      "postDate": "01/15/2016 01:08:18",
      "content": "<p>@Paul, this is a recurring question in various forms. </p>\n\n<p>The clinical reality is that the volumes are important and the EF is a derived number. So by making EF part of the submission score, you would tilt the competition towards methods that are good at predicting EF and so so at predicting volumes. As you point out, predicting EF is the easier problem and there are methods that predict EF pretty well, so it is less interesting because it is easier, and it would, in fact, be a less clinically useful solution. Specifically, a method that is able to predict EF cannot replace manual segmentation of the images. If you need to manually segment anyway, there is no need for a method that can predict EF. </p>\n\n<p>Hope this helps. </p>",
      "rawMarkdown": "Paul, this is a recurring question in various forms. \r\n\r\nThe clinical reality is that the volumes are important and the EF is a derived number. So by making EF part of the submission score, you would tilt the competition towards methods that are good at predicting EF and so so at predicting volumes. As you point out, predicting EF is the easier problem and there are methods that predict EF pretty well, so it is less interesting because it is easier, and it would, in fact, be a less clinically useful solution. Specifically, a method that is able to predict EF cannot replace manual segmentation of the images. If you need to manually segment anyway, there is no need for a method that can predict EF. \r\n\r\nHope this helps.",
      "votes": null
    },
    {
      "id": "104677",
      "postDate": "01/15/2016 06:17:25",
      "content": "<p>@Michael</p>\n\n<p>The solution I have in mind would be based entirely on segmentation - no gimmicks. The way it could be used clinically is to offer automated segmentation, which would be as good as manual for most LV slices, except of some top and bottom ones. For these areas one or more segmentation options could be presented to human operator to choose from or reject as belonging to the slice outside of LV. The net effect would be a dramatic reduction in manual labor, which I think is the objective of this competition. </p>",
      "rawMarkdown": "Michael\r\n\r\nThe solution I have in mind would be based entirely on segmentation - no gimmicks. The way it could be used clinically is to offer automated segmentation, which would be as good as manual for most LV slices, except of some top and bottom ones. For these areas one or more segmentation options could be presented to human operator to choose from or reject as belonging to the slice outside of LV. The net effect would be a dramatic reduction in manual labor, which I think is the objective of this competition.",
      "votes": null
    },
    {
      "id": "104687",
      "postDate": "01/15/2016 09:41:20",
      "content": "<p>@Paul the hard part of the segmentation is those basal slices where human intervention is needed a lot. There are existing algorithms that do a pretty reasonable job in automated segmentation of the easy slices. So the effect of what you are proposing would likely be less dramatic than you think. We wanted this competition to be directly aimed at the hard problem. I understand that it is a hard problem. </p>",
      "rawMarkdown": "Paul the hard part of the segmentation is those basal slices where human intervention is needed a lot. There are existing algorithms that do a pretty reasonable job in automated segmentation of the easy slices. So the effect of what you are proposing would likely be less dramatic than you think. We wanted this competition to be directly aimed at the hard problem. I understand that it is a hard problem.",
      "votes": null
    },
    {
      "id": "104688",
      "postDate": "01/15/2016 10:03:07",
      "content": "<p>@Michael</p>\n\n<p>Thank you for clarifying this issue. I have much better picture now of what you are aiming for.</p>",
      "rawMarkdown": "Michael\r\n\r\nThank you for clarifying this issue. I have much better picture now of what you are aiming for.",
      "votes": null
    },
    {
      "id": "111419",
      "postDate": "03/14/2016 16:32:24",
      "content": "<p>[quote=Michael Hansen;104687]</p>\n\n<p>@Paul the hard part of the segmentation is those basal slices where human intervention is needed a lot. There are existing algorithms that do a pretty reasonable job in automated segmentation of the easy slices. So the effect of what you are proposing would likely be less dramatic than you think. We wanted this competition to be directly aimed at the hard problem. I understand that it is a hard problem. </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "[quote=Michael Hansen;104687]\r\n\r\n@Paul the hard part of the segmentation is those basal slices where human intervention is needed a lot. There are existing algorithms that do a pretty reasonable job in automated segmentation of the easy slices. So the effect of what you are proposing would likely be less dramatic than you think. We wanted this competition to be directly aimed at the hard problem. I understand that it is a hard problem. \r\n\r\n[/quote]",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 104503,
      "author_name": "shannonlantzy",
      "author_url": "",
      "post_date": "01/13/2016 11:59:08",
      "content": "<p>Thanks to Jason Farbman at Booz Allen for putting together this graphic. (<a href=\"http://www.datasciencebowl.com/crps_and_its_implications/\">Re-posted from the Data Science Bowl Blog.</a>)</p>\n\n<p><img src=\"http://www.datasciencebowl.com/wp-content/uploads/2016/01/booz-dsb-err-3.png\" alt=\"CRPS and its implications\" title></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104646,
      "author_name": "woolsey",
      "author_url": "",
      "post_date": "01/14/2016 21:58:46",
      "content": "<p>CRPS is not very intuitive at first, but the more I explore it, the more I like it.</p>\n\n<p>Imho the key observation is that, once you made your best guesses, you'll score higher with a cumulative distribution that reflects the true variance instead of a step function. This forces competitors to reveal not only their best guesses, but also the amount of noise they believe would remain unexplained. </p>\n\n<p>Congrats admins for this very clever choice... fair and robust!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104655,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "01/14/2016 23:49:31",
      "content": "<p>@Shannon</p>\n\n<p>You've shown in your post how small volume measurement error can dramatically affect EF value. The problem I have with this example is that EF can be calculated much more accurately without relying on EDV and ESV. I think that you can obtain high precision measurement of LV contraction based only on clearly imaged slices and disregard the top and bottom ones, which are the most difficult to interpret and most likely contribute most of the volume measurement error.</p>\n\n<p>I wrote about it earlier, but I think that having EF as a part of submission score independent of EDV and ESV would lead to more clinically useful solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104659,
      "author_name": "michaelhansen",
      "author_url": "",
      "post_date": "01/15/2016 01:08:18",
      "content": "<p>@Paul, this is a recurring question in various forms. </p>\n\n<p>The clinical reality is that the volumes are important and the EF is a derived number. So by making EF part of the submission score, you would tilt the competition towards methods that are good at predicting EF and so so at predicting volumes. As you point out, predicting EF is the easier problem and there are methods that predict EF pretty well, so it is less interesting because it is easier, and it would, in fact, be a less clinically useful solution. Specifically, a method that is able to predict EF cannot replace manual segmentation of the images. If you need to manually segment anyway, there is no need for a method that can predict EF. </p>\n\n<p>Hope this helps. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104677,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "01/15/2016 06:17:25",
      "content": "<p>@Michael</p>\n\n<p>The solution I have in mind would be based entirely on segmentation - no gimmicks. The way it could be used clinically is to offer automated segmentation, which would be as good as manual for most LV slices, except of some top and bottom ones. For these areas one or more segmentation options could be presented to human operator to choose from or reject as belonging to the slice outside of LV. The net effect would be a dramatic reduction in manual labor, which I think is the objective of this competition. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104687,
      "author_name": "michaelhansen",
      "author_url": "",
      "post_date": "01/15/2016 09:41:20",
      "content": "<p>@Paul the hard part of the segmentation is those basal slices where human intervention is needed a lot. There are existing algorithms that do a pretty reasonable job in automated segmentation of the easy slices. So the effect of what you are proposing would likely be less dramatic than you think. We wanted this competition to be directly aimed at the hard problem. I understand that it is a hard problem. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104688,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "01/15/2016 10:03:07",
      "content": "<p>@Michael</p>\n\n<p>Thank you for clarifying this issue. I have much better picture now of what you are aiming for.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111419,
      "author_name": "srikanthvvgs",
      "author_url": "",
      "post_date": "03/14/2016 16:32:24",
      "content": "<p>[quote=Michael Hansen;104687]</p>\n\n<p>@Paul the hard part of the segmentation is those basal slices where human intervention is needed a lot. There are existing algorithms that do a pretty reasonable job in automated segmentation of the easy slices. So the effect of what you are proposing would likely be less dramatic than you think. We wanted this competition to be directly aimed at the hard problem. I understand that it is a hard problem. </p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "101878": "We've received some inquiries/concerns about the \"low\" early competition scores, enough to warrant a thread about what it means for the CRPS to be \"low\".\r\n\r\nWith any metric (but especially mathematically complex ones) it's important not to let intuition overwhelm your interpretation. It's easy to *feel* that 1 is high and 0 is low, and therefore the current scores are in danger of saturation. However, you should keep in mind three important points here:\r\n\r\n 1. The exponent in the metric is squaring a number < 1 (yielding an even smaller number).\r\n 2. You can and should plot your probability curves on top of a vertical line for the volume. If you do this, you can visually see that there is a lot of room for improvement, even for CRPS scores that feel low.\r\n 3. Medical use cases impose a very high bar on the level of robustness and accuracy for an algorithm to make it into the clinic.\r\n\r\nWhat's the takeaway? When you look at the leaderboard and see a number that feels low, or gaps that seem very close, view the plots and run the numbers before you get discouraged. It's a loooong road to get to 0.0!",
    "104503": "Thanks to Jason Farbman at Booz Allen for putting together this graphic. ([Re-posted from the Data Science Bowl Blog.][2])\r\n\r\n\r\n ![CRPS and its implications][1]\r\n\r\n\r\n\r\n  [1]: http://www.datasciencebowl.com/wp-content/uploads/2016/01/booz-dsb-err-3.png\r\n  [2]: http://www.datasciencebowl.com/crps_and_its_implications/",
    "104646": "CRPS is not very intuitive at first, but the more I explore it, the more I like it.\r\n\r\nImho the key observation is that, once you made your best guesses, you'll score higher with a cumulative distribution that reflects the true variance instead of a step function. This forces competitors to reveal not only their best guesses, but also the amount of noise they believe would remain unexplained. \r\n\r\nCongrats admins for this very clever choice... fair and robust!",
    "104655": "Shannon\r\n\r\nYou've shown in your post how small volume measurement error can dramatically affect EF value. The problem I have with this example is that EF can be calculated much more accurately without relying on EDV and ESV. I think that you can obtain high precision measurement of LV contraction based only on clearly imaged slices and disregard the top and bottom ones, which are the most difficult to interpret and most likely contribute most of the volume measurement error.\r\n\r\nI wrote about it earlier, but I think that having EF as a part of submission score independent of EDV and ESV would lead to more clinically useful solution.",
    "104659": "Paul, this is a recurring question in various forms. \r\n\r\nThe clinical reality is that the volumes are important and the EF is a derived number. So by making EF part of the submission score, you would tilt the competition towards methods that are good at predicting EF and so so at predicting volumes. As you point out, predicting EF is the easier problem and there are methods that predict EF pretty well, so it is less interesting because it is easier, and it would, in fact, be a less clinically useful solution. Specifically, a method that is able to predict EF cannot replace manual segmentation of the images. If you need to manually segment anyway, there is no need for a method that can predict EF. \r\n\r\nHope this helps.",
    "104677": "Michael\r\n\r\nThe solution I have in mind would be based entirely on segmentation - no gimmicks. The way it could be used clinically is to offer automated segmentation, which would be as good as manual for most LV slices, except of some top and bottom ones. For these areas one or more segmentation options could be presented to human operator to choose from or reject as belonging to the slice outside of LV. The net effect would be a dramatic reduction in manual labor, which I think is the objective of this competition.",
    "104687": "Paul the hard part of the segmentation is those basal slices where human intervention is needed a lot. There are existing algorithms that do a pretty reasonable job in automated segmentation of the easy slices. So the effect of what you are proposing would likely be less dramatic than you think. We wanted this competition to be directly aimed at the hard problem. I understand that it is a hard problem.",
    "104688": "Michael\r\n\r\nThank you for clarifying this issue. I have much better picture now of what you are aiming for.",
    "111419": "[quote=Michael Hansen;104687]\r\n\r\n@Paul the hard part of the segmentation is those basal slices where human intervention is needed a lot. There are existing algorithms that do a pretty reasonable job in automated segmentation of the easy slices. So the effect of what you are proposing would likely be less dramatic than you think. We wanted this competition to be directly aimed at the hard problem. I understand that it is a hard problem. \r\n\r\n[/quote]"
  },
  "source": "meta"
}