{
  "id": 18116,
  "title": "Evaluation Metric",
  "url": "/competitions/second-annual-data-science-bowl/discussion/18116",
  "author_name": "",
  "post_date": "2015-12-24T21:17:00.933Z",
  "votes": 4,
  "comment_count": 3,
  "views": 1130,
  "content": "<p>I wanted to relate back the evaluation metric to something a little more comprehensible, to get an idea of what certain values indicate in a crude simplification of a model.</p>\n\n<p>Imagine we make a model which always outputs uniform CDFs within +/- (delta) mL of the target 100% of the time, centered at the actual target. For example, when the target is 300 mL, and we have a model with an implied delta of 1 mL, imagine our model outputing:</p>\n\n<p>297 , 298 , 299 , 300 , 301 , 302 , 303 </p>\n\n<p>0 , 0 , .25 , .50 , .75 , 1 , 1 </p>\n\n<p>Obviously this is a huge oversimplification / generalization of a model output (and would be one HECK of a model!), but we can do some quick math to relate back our evaluation metric to certain values of delta.</p>\n\n<p>Note the log scale on both axes <a href=\"http://imgur.com/49EZb0K\">relationship</a></p>\n\n<p>So we can think of the leader right now (.02866) as essentially (though certainly not exactly) being able to estimate the target to within +/- 100 mL, which is a HUGE improvement on always outputing the CDF of the training set, which is equivalent to +/- 150 mL.</p>\n\n<p>If we imagine what might be &quot;acceptable&quot; in a clinical setting, say, +/- 10 mL, we're looking at roughly an evaluation score of .003.</p>\n\n<p>Again, this is an oversimplification of how an imaginary model could work, but hopefully this will help with anyone trying to relate back the evaluation metric to something a little more understandable.</p>",
  "messages": [
    {
      "id": "102709",
      "postDate": "12/24/2015 21:17:00",
      "content": "<p>I wanted to relate back the evaluation metric to something a little more comprehensible, to get an idea of what certain values indicate in a crude simplification of a model.</p>\n\n<p>Imagine we make a model which always outputs uniform CDFs within +/- (delta) mL of the target 100% of the time, centered at the actual target. For example, when the target is 300 mL, and we have a model with an implied delta of 1 mL, imagine our model outputing:</p>\n\n<p>297 , 298 , 299 , 300 , 301 , 302 , 303 </p>\n\n<p>0 , 0 , .25 , .50 , .75 , 1 , 1 </p>\n\n<p>Obviously this is a huge oversimplification / generalization of a model output (and would be one HECK of a model!), but we can do some quick math to relate back our evaluation metric to certain values of delta.</p>\n\n<p>Note the log scale on both axes <a href=\"http://imgur.com/49EZb0K\">relationship</a></p>\n\n<p>So we can think of the leader right now (.02866) as essentially (though certainly not exactly) being able to estimate the target to within +/- 100 mL, which is a HUGE improvement on always outputing the CDF of the training set, which is equivalent to +/- 150 mL.</p>\n\n<p>If we imagine what might be &quot;acceptable&quot; in a clinical setting, say, +/- 10 mL, we're looking at roughly an evaluation score of .003.</p>\n\n<p>Again, this is an oversimplification of how an imaginary model could work, but hopefully this will help with anyone trying to relate back the evaluation metric to something a little more understandable.</p>",
      "rawMarkdown": "I wanted to relate back the evaluation metric to something a little more comprehensible, to get an idea of what certain values indicate in a crude simplification of a model.\r\n\r\nImagine we make a model which always outputs uniform CDFs within +/- (delta) mL of the target 100% of the time, centered at the actual target. For example, when the target is 300 mL, and we have a model with an implied delta of 1 mL, imagine our model outputing:\r\n\r\n297 , 298 , 299 , 300 , 301 , 302 , 303 \r\n\r\n0 , 0 , .25 , .50 , .75 , 1 , 1 \r\n\r\nObviously this is a huge oversimplification / generalization of a model output (and would be one HECK of a model!), but we can do some quick math to relate back our evaluation metric to certain values of delta.\r\n\r\nNote the log scale on both axes [relationship][1]\r\n\r\nSo we can think of the leader right now (.02866) as essentially (though certainly not exactly) being able to estimate the target to within +/- 100 mL, which is a HUGE improvement on always outputing the CDF of the training set, which is equivalent to +/- 150 mL.\r\n\r\nIf we imagine what might be \"acceptable\" in a clinical setting, say, +/- 10 mL, we're looking at roughly an evaluation score of .003.\r\n\r\nAgain, this is an oversimplification of how an imaginary model could work, but hopefully this will help with anyone trying to relate back the evaluation metric to something a little more understandable.\r\n\r\n  [1]: http://imgur.com/49EZb0K",
      "votes": null
    },
    {
      "id": "103269",
      "postDate": "12/31/2015 02:42:07",
      "content": "<p>Thanks for the description, Martin.  As long as we're along the lines of talking about the evaluation metric, does anyone else think it odd that we're using a metric that equally weights the absolute errors of the systole and the diastole?  Given that the diastole is typically larger, it seems that it will have the smaller relative error compared to that on the systole for the same level of absolute error.  In this case, it would seem that the relative error would be a bit more pertinent.</p>",
      "rawMarkdown": "Thanks for the description, Martin.  As long as we're along the lines of talking about the evaluation metric, does anyone else think it odd that we're using a metric that equally weights the absolute errors of the systole and the diastole?  Given that the diastole is typically larger, it seems that it will have the smaller relative error compared to that on the systole for the same level of absolute error.  In this case, it would seem that the relative error would be a bit more pertinent.",
      "votes": null
    },
    {
      "id": "103296",
      "postDate": "12/31/2015 14:39:07",
      "content": "<p>The generic formula would be delta/3600.</p>",
      "rawMarkdown": "The generic formula would be delta/3600.",
      "votes": null
    },
    {
      "id": "103494",
      "postDate": "01/03/2016 17:45:21",
      "content": "<p>I think the results also depend on the correct code for the CDF. It could be fine to compare CDF codes (matlab/python) to assure yourself too if you are getting a correct algorithm or not.</p>",
      "rawMarkdown": "I think the results also depend on the correct code for the CDF. It could be fine to compare CDF codes (matlab/python) to assure yourself too if you are getting a correct algorithm or not.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 103269,
      "author_name": "csgwon",
      "author_url": "",
      "post_date": "12/31/2015 02:42:07",
      "content": "<p>Thanks for the description, Martin.  As long as we're along the lines of talking about the evaluation metric, does anyone else think it odd that we're using a metric that equally weights the absolute errors of the systole and the diastole?  Given that the diastole is typically larger, it seems that it will have the smaller relative error compared to that on the systole for the same level of absolute error.  In this case, it would seem that the relative error would be a bit more pertinent.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103296,
      "author_name": "udayabhanu",
      "author_url": "",
      "post_date": "12/31/2015 14:39:07",
      "content": "<p>The generic formula would be delta/3600.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103494,
      "author_name": "liveflow",
      "author_url": "",
      "post_date": "01/03/2016 17:45:21",
      "content": "<p>I think the results also depend on the correct code for the CDF. It could be fine to compare CDF codes (matlab/python) to assure yourself too if you are getting a correct algorithm or not.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "102709": "I wanted to relate back the evaluation metric to something a little more comprehensible, to get an idea of what certain values indicate in a crude simplification of a model.\r\n\r\nImagine we make a model which always outputs uniform CDFs within +/- (delta) mL of the target 100% of the time, centered at the actual target. For example, when the target is 300 mL, and we have a model with an implied delta of 1 mL, imagine our model outputing:\r\n\r\n297 , 298 , 299 , 300 , 301 , 302 , 303 \r\n\r\n0 , 0 , .25 , .50 , .75 , 1 , 1 \r\n\r\nObviously this is a huge oversimplification / generalization of a model output (and would be one HECK of a model!), but we can do some quick math to relate back our evaluation metric to certain values of delta.\r\n\r\nNote the log scale on both axes [relationship][1]\r\n\r\nSo we can think of the leader right now (.02866) as essentially (though certainly not exactly) being able to estimate the target to within +/- 100 mL, which is a HUGE improvement on always outputing the CDF of the training set, which is equivalent to +/- 150 mL.\r\n\r\nIf we imagine what might be \"acceptable\" in a clinical setting, say, +/- 10 mL, we're looking at roughly an evaluation score of .003.\r\n\r\nAgain, this is an oversimplification of how an imaginary model could work, but hopefully this will help with anyone trying to relate back the evaluation metric to something a little more understandable.\r\n\r\n  [1]: http://imgur.com/49EZb0K",
    "103269": "Thanks for the description, Martin.  As long as we're along the lines of talking about the evaluation metric, does anyone else think it odd that we're using a metric that equally weights the absolute errors of the systole and the diastole?  Given that the diastole is typically larger, it seems that it will have the smaller relative error compared to that on the systole for the same level of absolute error.  In this case, it would seem that the relative error would be a bit more pertinent.",
    "103296": "The generic formula would be delta/3600.",
    "103494": "I think the results also depend on the correct code for the CDF. It could be fine to compare CDF codes (matlab/python) to assure yourself too if you are getting a correct algorithm or not."
  },
  "source": "meta"
}