{
  "id": 17926,
  "title": "Problems with evaluation description",
  "url": "/competitions/second-annual-data-science-bowl/discussion/17926",
  "author_name": "",
  "post_date": "2015-12-15T16:13:15.887Z",
  "votes": 6,
  "comment_count": 4,
  "views": 1184,
  "content": "<p>Thank you very much for hosting this interesting competition!</p>\n\n<p>I think there is a small problem with the <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/details/evaluation\">description of the evaluation measure</a>:</p>\n\n<p>Firstly, it might be clearer if the formula was written as $$C=\\frac{1}{600N}\\sum_{m=1}^{N}\\sum_{n=0}^{599} \\left(P(y\\leq n)-H(n-V_m)\\right)^2\\;,$$ i.e., explicitly writing the sum over \\(m\\). </p>\n\n<p>Secondly, this score is <em>not</em> equal to the area in the given figure. This would be the case if the summand was \\(\\left|P(y\\leq n)-H(n-V_m)\\right|\\). Due to the square it is rather the volume of the shape you obtain when rotating the two areas around \\(y=0\\) and \\(y=1\\), respectively. This is not easy to visualize such that it might make sense to drop the (slightly misleading) figure altogether.</p>\n\n<p>Cheers,\nML</p>",
  "messages": [
    {
      "id": "101495",
      "postDate": "12/15/2015 16:13:15",
      "content": "<p>Thank you very much for hosting this interesting competition!</p>\n\n<p>I think there is a small problem with the <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/details/evaluation\">description of the evaluation measure</a>:</p>\n\n<p>Firstly, it might be clearer if the formula was written as $$C=\\frac{1}{600N}\\sum_{m=1}^{N}\\sum_{n=0}^{599} \\left(P(y\\leq n)-H(n-V_m)\\right)^2\\;,$$ i.e., explicitly writing the sum over \\(m\\). </p>\n\n<p>Secondly, this score is <em>not</em> equal to the area in the given figure. This would be the case if the summand was \\(\\left|P(y\\leq n)-H(n-V_m)\\right|\\). Due to the square it is rather the volume of the shape you obtain when rotating the two areas around \\(y=0\\) and \\(y=1\\), respectively. This is not easy to visualize such that it might make sense to drop the (slightly misleading) figure altogether.</p>\n\n<p>Cheers,\nML</p>",
      "rawMarkdown": "Thank you very much for hosting this interesting competition!\r\n\r\nI think there is a small problem with the [description of the evaluation measure][1]:\r\n\r\nFirstly, it might be clearer if the formula was written as $$C=\\frac{1}{600N}\\sum_{m=1}^{N}\\sum_{n=0}^{599} \\left(P(y\\leq n)-H(n-V_m)\\right)^2\\;,$$ i.e., explicitly writing the sum over \\\\(m\\\\). \r\n\r\nSecondly, this score is *not* equal to the area in the given figure. This would be the case if the summand was \\\\(\\left|P(y\\leq n)-H(n-V_m)\\right|\\\\). Due to the square it is rather the volume of the shape you obtain when rotating the two areas around \\\\(y=0\\\\) and \\\\(y=1\\\\), respectively. This is not easy to visualize such that it might make sense to drop the (slightly misleading) figure altogether.\r\n\r\nCheers,\r\nML\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/second-annual-data-science-bowl/details/evaluation",
      "votes": null
    },
    {
      "id": "101499",
      "postDate": "12/15/2015 16:23:25",
      "content": "<p>Appreciate the feedback and you're absolutely correct. I've updated the formula and changed the wording pertaining to the figure. Even though it's not exact, I think the figure is a valuable aid for understanding what the metric is doing.</p>",
      "rawMarkdown": "Appreciate the feedback and you're absolutely correct. I've updated the formula and changed the wording pertaining to the figure. Even though it's not exact, I think the figure is a valuable aid for understanding what the metric is doing.",
      "votes": null
    },
    {
      "id": "101511",
      "postDate": "12/15/2015 17:32:33",
      "content": "<p>Thank you very much for the quick reply and fix.</p>",
      "rawMarkdown": "Thank you very much for the quick reply and fix.",
      "votes": null
    },
    {
      "id": "104287",
      "postDate": "01/11/2016 12:37:29",
      "content": "<p>There is a big difference in evaluation results between using an absolute value vs. square in these sums. </p>\n\n<p>In case of absolute values, which yield score equal to the area on the figure from <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/details/evaluation\">Evaluation</a> chapter, there is no need to optimize the shape of predicted probability distribution curve when one can predict an exact volume. It seems to me, at least for curves symmetric about <strong><em>y = 0.5</em></strong>, that a simple step function going from 0 to 1 at predicted volume is optimal.</p>\n\n<p>This is not a case for squared differences currently used in evaluation formula. Step function is far from optimal, at least in some cases. I just jumped 109 positions in leaderboard simply by replacing vertical step with a slope in my probability distribution curves. This is somewhat distracting from the main objective, so I wonder if anyone knows the optimal shape to replace step function with. </p>",
      "rawMarkdown": "There is a big difference in evaluation results between using an absolute value vs. square in these sums. \r\n\r\nIn case of absolute values, which yield score equal to the area on the figure from [Evaluation][1] chapter, there is no need to optimize the shape of predicted probability distribution curve when one can predict an exact volume. It seems to me, at least for curves symmetric about ***y = 0.5***, that a simple step function going from 0 to 1 at predicted volume is optimal.\r\n\r\nThis is not a case for squared differences currently used in evaluation formula. Step function is far from optimal, at least in some cases. I just jumped 109 positions in leaderboard simply by replacing vertical step with a slope in my probability distribution curves. This is somewhat distracting from the main objective, so I wonder if anyone knows the optimal shape to replace step function with. \r\n\r\n\r\n  [1]: https://www.kaggle.com/c/second-annual-data-science-bowl/details/evaluation",
      "votes": null
    },
    {
      "id": "104458",
      "postDate": "01/12/2016 21:52:14",
      "content": "<p>@Paul Jurczak, I said something about this in another <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18352/clinical-significance\">thread</a>, but since then I've done some experiments and it appears that using:\n$$\n\\text{CDF}(v) = \\frac{1}{2}\\text{erfc}\\left(\\frac{v_0 - v}{\\sqrt{2}\\sigma} \\right)\n$$\nworks quite well. \\(v_0\\) is the prediction of the model and sigma should be chosen to represent the uncertainty of the model. If your model doesn't put out any sort of estimate of uncertainty, you can just treat it as a parameter and chose it using cross validation.  So far I've only looked at this using local CV, but that's been pretty reliable for me so far in this contest.</p>\n\n<p>Hope that's useful.</p>\n\n<p>[EDIT] OK, I used the above formulation as part of what was supposed to be my new improved post processing chain. It turned out not to be an improvement, but neither was it much of a regression. Predicting the CDF directly gives me 0.014945 on the leaderboard while running it through my post processing chain, which includes converting too and from \\((v_0, \\sigma)\\) representation, gives me 0.015401. So, it's perfectly possible to generate a respectable score using this representation.</p>",
      "rawMarkdown": "Paul Jurczak, I said something about this in another [thread][1], but since then I've done some experiments and it appears that using:\r\n$$\r\n\\text{CDF}(v) = \\frac{1}{2}\\text{erfc}\\left(\\frac{v_0 - v}{\\sqrt{2}\\sigma} \\right)\r\n$$\r\nworks quite well. \\\\(v_0\\\\) is the prediction of the model and sigma should be chosen to represent the uncertainty of the model. If your model doesn't put out any sort of estimate of uncertainty, you can just treat it as a parameter and chose it using cross validation.  So far I've only looked at this using local CV, but that's been pretty reliable for me so far in this contest.\r\n\r\nHope that's useful.\r\n\r\n[EDIT] OK, I used the above formulation as part of what was supposed to be my new improved post processing chain. It turned out not to be an improvement, but neither was it much of a regression. Predicting the CDF directly gives me 0.014945 on the leaderboard while running it through my post processing chain, which includes converting too and from \\\\((v_0, \\sigma)\\\\) representation, gives me 0.015401. So, it's perfectly possible to generate a respectable score using this representation.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18352/clinical-significance",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 101499,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "12/15/2015 16:23:25",
      "content": "<p>Appreciate the feedback and you're absolutely correct. I've updated the formula and changed the wording pertaining to the figure. Even though it's not exact, I think the figure is a valuable aid for understanding what the metric is doing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101511,
      "author_name": "mlegner",
      "author_url": "",
      "post_date": "12/15/2015 17:32:33",
      "content": "<p>Thank you very much for the quick reply and fix.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104287,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "01/11/2016 12:37:29",
      "content": "<p>There is a big difference in evaluation results between using an absolute value vs. square in these sums. </p>\n\n<p>In case of absolute values, which yield score equal to the area on the figure from <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/details/evaluation\">Evaluation</a> chapter, there is no need to optimize the shape of predicted probability distribution curve when one can predict an exact volume. It seems to me, at least for curves symmetric about <strong><em>y = 0.5</em></strong>, that a simple step function going from 0 to 1 at predicted volume is optimal.</p>\n\n<p>This is not a case for squared differences currently used in evaluation formula. Step function is far from optimal, at least in some cases. I just jumped 109 positions in leaderboard simply by replacing vertical step with a slope in my probability distribution curves. This is somewhat distracting from the main objective, so I wonder if anyone knows the optimal shape to replace step function with. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104458,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "01/12/2016 21:52:14",
      "content": "<p>@Paul Jurczak, I said something about this in another <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18352/clinical-significance\">thread</a>, but since then I've done some experiments and it appears that using:\n$$\n\\text{CDF}(v) = \\frac{1}{2}\\text{erfc}\\left(\\frac{v_0 - v}{\\sqrt{2}\\sigma} \\right)\n$$\nworks quite well. \\(v_0\\) is the prediction of the model and sigma should be chosen to represent the uncertainty of the model. If your model doesn't put out any sort of estimate of uncertainty, you can just treat it as a parameter and chose it using cross validation.  So far I've only looked at this using local CV, but that's been pretty reliable for me so far in this contest.</p>\n\n<p>Hope that's useful.</p>\n\n<p>[EDIT] OK, I used the above formulation as part of what was supposed to be my new improved post processing chain. It turned out not to be an improvement, but neither was it much of a regression. Predicting the CDF directly gives me 0.014945 on the leaderboard while running it through my post processing chain, which includes converting too and from \\((v_0, \\sigma)\\) representation, gives me 0.015401. So, it's perfectly possible to generate a respectable score using this representation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "101495": "Thank you very much for hosting this interesting competition!\r\n\r\nI think there is a small problem with the [description of the evaluation measure][1]:\r\n\r\nFirstly, it might be clearer if the formula was written as $$C=\\frac{1}{600N}\\sum_{m=1}^{N}\\sum_{n=0}^{599} \\left(P(y\\leq n)-H(n-V_m)\\right)^2\\;,$$ i.e., explicitly writing the sum over \\\\(m\\\\). \r\n\r\nSecondly, this score is *not* equal to the area in the given figure. This would be the case if the summand was \\\\(\\left|P(y\\leq n)-H(n-V_m)\\right|\\\\). Due to the square it is rather the volume of the shape you obtain when rotating the two areas around \\\\(y=0\\\\) and \\\\(y=1\\\\), respectively. This is not easy to visualize such that it might make sense to drop the (slightly misleading) figure altogether.\r\n\r\nCheers,\r\nML\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/second-annual-data-science-bowl/details/evaluation",
    "101499": "Appreciate the feedback and you're absolutely correct. I've updated the formula and changed the wording pertaining to the figure. Even though it's not exact, I think the figure is a valuable aid for understanding what the metric is doing.",
    "101511": "Thank you very much for the quick reply and fix.",
    "104287": "There is a big difference in evaluation results between using an absolute value vs. square in these sums. \r\n\r\nIn case of absolute values, which yield score equal to the area on the figure from [Evaluation][1] chapter, there is no need to optimize the shape of predicted probability distribution curve when one can predict an exact volume. It seems to me, at least for curves symmetric about ***y = 0.5***, that a simple step function going from 0 to 1 at predicted volume is optimal.\r\n\r\nThis is not a case for squared differences currently used in evaluation formula. Step function is far from optimal, at least in some cases. I just jumped 109 positions in leaderboard simply by replacing vertical step with a slope in my probability distribution curves. This is somewhat distracting from the main objective, so I wonder if anyone knows the optimal shape to replace step function with. \r\n\r\n\r\n  [1]: https://www.kaggle.com/c/second-annual-data-science-bowl/details/evaluation",
    "104458": "Paul Jurczak, I said something about this in another [thread][1], but since then I've done some experiments and it appears that using:\r\n$$\r\n\\text{CDF}(v) = \\frac{1}{2}\\text{erfc}\\left(\\frac{v_0 - v}{\\sqrt{2}\\sigma} \\right)\r\n$$\r\nworks quite well. \\\\(v_0\\\\) is the prediction of the model and sigma should be chosen to represent the uncertainty of the model. If your model doesn't put out any sort of estimate of uncertainty, you can just treat it as a parameter and chose it using cross validation.  So far I've only looked at this using local CV, but that's been pretty reliable for me so far in this contest.\r\n\r\nHope that's useful.\r\n\r\n[EDIT] OK, I used the above formulation as part of what was supposed to be my new improved post processing chain. It turned out not to be an improvement, but neither was it much of a regression. Predicting the CDF directly gives me 0.014945 on the leaderboard while running it through my post processing chain, which includes converting too and from \\\\((v_0, \\sigma)\\\\) representation, gives me 0.015401. So, it's perfectly possible to generate a respectable score using this representation.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18352/clinical-significance"
  },
  "source": "meta"
}