{
  "id": 17966,
  "title": "Optimal submission for Continuous Ranked Probability Score (CRPS)",
  "url": "/competitions/second-annual-data-science-bowl/discussion/17966",
  "author_name": "",
  "post_date": "2015-12-16T23:26:09.967Z",
  "votes": 9,
  "comment_count": 12,
  "views": 2680,
  "content": "<p>short math on why 0.042738 is the best score you can get without looking at the images\n<a href=\"https://github.com/udibr/DSB2/blob/master/optimal.ipynb\">link</a></p>",
  "messages": [
    {
      "id": "101765",
      "postDate": "12/16/2015 23:26:09",
      "content": "<p>short math on why 0.042738 is the best score you can get without looking at the images\n<a href=\"https://github.com/udibr/DSB2/blob/master/optimal.ipynb\">link</a></p>",
      "rawMarkdown": "short math on why 0.042738 is the best score you can get without looking at the images\r\n[link][1]\r\n \r\n\r\n\r\n  [1]: https://github.com/udibr/DSB2/blob/master/optimal.ipynb",
      "votes": null
    },
    {
      "id": "101931",
      "postDate": "12/17/2015 23:08:56",
      "content": "<p>Thank you!</p>\n\n<p>Can you show more details how do you take the derivative to get the following equation?</p>\n\n<p>${ \\partial C \\over \\partial q_w } = E_{p_v} \\left( \\sum_n 2 (Q_n -H(n \\ge v )) H(n \\ge w)\\right) + \\lambda </p>",
      "rawMarkdown": "Thank you!\r\n\r\nCan you show more details how do you take the derivative to get the following equation?\r\n\r\n${ \\partial C \\over \\partial q_w } = E_{p_v} \\left( \\sum_n 2 (Q_n -H(n \\ge v )) H(n \\ge w)\\right) + \\lambda",
      "votes": null
    },
    {
      "id": "101932",
      "postDate": "12/17/2015 23:23:51",
      "content": "<p>Each q_j is a different variable and you have 600 of them. So you are actually making 600 different derivatives.  The part with the \\lambda in front of it is easy. In the sum each q_j appears once and when you decide to take a derivative according to just one of the q (for example q_w) then the derivative of each is part of the sum is zero unless you happen to take the derivative of q_w and in that case it is one. So the derivative of the entire sum is always one regardless of w.</p>\n\n<p>The first part is a little bit more complex. When you take the derivative of  (Q_n -H(n \\ge v ))^2 you have to use the chain rule. First you work the ^2 which gives you 2 (Q_n -H(n \\ge v )) but then you have to take the derivative of Q_n with respect to q_w. </p>\n\n<p>Q_n is just the sum of q_j up to n. So if n is \\ge w then q_w will appear in this sum and the derivative will be  one. But if n is \\lt w then q_w will be absent and the derivative will be zero. Therefore the derivative of Q_n is just H(n \\ge w)</p>\n\n<p>HTH, Udi</p>",
      "rawMarkdown": "Each q_j is a different variable and you have 600 of them. So you are actually making 600 different derivatives.  The part with the \\lambda in front of it is easy. In the sum each q_j appears once and when you decide to take a derivative according to just one of the q (for example q_w) then the derivative of each is part of the sum is zero unless you happen to take the derivative of q_w and in that case it is one. So the derivative of the entire sum is always one regardless of w.\r\n\r\nThe first part is a little bit more complex. When you take the derivative of  (Q_n -H(n \\ge v ))^2 you have to use the chain rule. First you work the ^2 which gives you 2 (Q_n -H(n \\ge v )) but then you have to take the derivative of Q_n with respect to q_w. \r\n\r\nQ_n is just the sum of q_j up to n. So if n is \\ge w then q_w will appear in this sum and the derivative will be  one. But if n is \\lt w then q_w will be absent and the derivative will be zero. Therefore the derivative of Q_n is just H(n \\ge w)\r\n\r\nHTH, Udi",
      "votes": null
    },
    {
      "id": "101939",
      "postDate": "12/18/2015 00:23:39",
      "content": "<p>That makes sense! Thank you very much!</p>",
      "rawMarkdown": "That makes sense! Thank you very much!",
      "votes": null
    },
    {
      "id": "101997",
      "postDate": "12/18/2015 13:02:54",
      "content": "<p>I guess javascript is disabled on your browser. You can read the document in text form at <a href=\"https://raw.githubusercontent.com/udibr/DSB2/master/optimal.ipynb\">https://raw.githubusercontent.com/udibr/DSB2/master/optimal.ipynb</a></p>",
      "rawMarkdown": "I guess javascript is disabled on your browser. You can read the document in text form at https://raw.githubusercontent.com/udibr/DSB2/master/optimal.ipynb",
      "votes": null
    },
    {
      "id": "102001",
      "postDate": "12/18/2015 14:12:28",
      "content": "<p>[quote=udibr;101765]</p>\n\n<p>short math on why 0.042738 is the best score you can get without looking at the images\n<a href=\"https://github.com/udibr/DSB2/blob/master/optimal.ipynb\">link</a></p>\n\n<p>[/quote]</p>\n\n<p>Thank you udibr.\nI agree with you that the best score you can get without looking at the images is by using the cdf.\nBut the true cdf is not known and you use an empirical estimation in your code. So in my opinion and mathematically speaking, 0.042738 is not the best score because others estimations of the cdf exist. Anyway the best score is most likely very close to 0.042738.</p>",
      "rawMarkdown": "[quote=udibr;101765]\r\n\r\nshort math on why 0.042738 is the best score you can get without looking at the images\r\n[link][1]\r\n \r\n\r\n\r\n  [1]: https://github.com/udibr/DSB2/blob/master/optimal.ipynb\r\n\r\n[/quote]\r\n\r\nThank you udibr.\r\nI agree with you that the best score you can get without looking at the images is by using the cdf.\r\nBut the true cdf is not known and you use an empirical estimation in your code. So in my opinion and mathematically speaking, 0.042738 is not the best score because others estimations of the cdf exist. Anyway the best score is most likely very close to 0.042738.",
      "votes": null
    },
    {
      "id": "102015",
      "postDate": "12/18/2015 16:49:37",
      "content": "<p>I agree. You can also overfit the validation set and reach even lower score.</p>",
      "rawMarkdown": "I agree. You can also overfit the validation set and reach even lower score.",
      "votes": null
    },
    {
      "id": "102057",
      "postDate": "12/18/2015 23:47:22",
      "content": "<p>Hmm, assume the true distribution is a uniform distribution -- 1/600 for each volume bucket. Then, wouldn't the optimal guess for the CDF be 0 for 0 &lt; x &lt; 300, and 1 for x &gt;= 300, to get an expected score of 0.25 per sample? If you were to guess the true linear CDF, it seems you would get 0.33 per sample.</p>",
      "rawMarkdown": "Hmm, assume the true distribution is a uniform distribution -- 1/600 for each volume bucket. Then, wouldn't the optimal guess for the CDF be 0 for 0 < x < 300, and 1 for x >= 300, to get an expected score of 0.25 per sample? If you were to guess the true linear CDF, it seems you would get 0.33 per sample.",
      "votes": null
    },
    {
      "id": "102065",
      "postDate": "12/19/2015 01:44:44",
      "content": "<p>for a step CDF you will get 0.5*(1/600)^2 + 0.5*(1-1/600)^2 = 0.499</p>",
      "rawMarkdown": "for a step CDF you will get 0.5*(1/600)^2 + 0.5*(1-1/600)^2 = 0.499",
      "votes": null
    },
    {
      "id": "102066",
      "postDate": "12/19/2015 01:46:24",
      "content": "<p>Oh derp, I didn't see the square term. My bad!</p>",
      "rawMarkdown": "Oh derp, I didn't see the square term. My bad!",
      "votes": null
    },
    {
      "id": "103954",
      "postDate": "01/08/2016 06:49:05",
      "content": "<p>Udibr,  shouldn't this be &gt;= instead of =? </p>\n\n<p>\\sum_{n=w} Q_n</p>\n\n<p>And, please explain to me how did you get 0.499 for step CDF, what is x there? Thanks.</p>",
      "rawMarkdown": "Udibr,  shouldn't this be >= instead of =? \r\n\r\n\\\\sum_{n=w} Q_n\r\n\r\nAnd, please explain to me how did you get 0.499 for step CDF, what is x there? Thanks.",
      "votes": null
    },
    {
      "id": "104081",
      "postDate": "01/09/2016 00:13:54",
      "content": "<p>I didn't write the upward limit of the sum, M. The sum starts from n=w and runs upwards until M (so &gt;= would have been another way of saying it)\nI'm not sure I understand the second part of your question</p>",
      "rawMarkdown": "I didn't write the upward limit of the sum, M. The sum starts from n=w and runs upwards until M (so >= would have been another way of saying it)\r\nI'm not sure I understand the second part of your question",
      "votes": null
    },
    {
      "id": "104289",
      "postDate": "01/11/2016 12:55:59",
      "content": "<p>Well, I'm already at 0.036361 without doing any image processing or machine learning. It is a gimmick, but it works well for testing my software infrastructure. I will improve it a little bit and make a post about it.</p>",
      "rawMarkdown": "Well, I'm already at 0.036361 without doing any image processing or machine learning. It is a gimmick, but it works well for testing my software infrastructure. I will improve it a little bit and make a post about it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 101931,
      "author_name": "nkhuyu",
      "author_url": "",
      "post_date": "12/17/2015 23:08:56",
      "content": "<p>Thank you!</p>\n\n<p>Can you show more details how do you take the derivative to get the following equation?</p>\n\n<p>${ \\partial C \\over \\partial q_w } = E_{p_v} \\left( \\sum_n 2 (Q_n -H(n \\ge v )) H(n \\ge w)\\right) + \\lambda </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101932,
      "author_name": "udibr1",
      "author_url": "",
      "post_date": "12/17/2015 23:23:51",
      "content": "<p>Each q_j is a different variable and you have 600 of them. So you are actually making 600 different derivatives.  The part with the \\lambda in front of it is easy. In the sum each q_j appears once and when you decide to take a derivative according to just one of the q (for example q_w) then the derivative of each is part of the sum is zero unless you happen to take the derivative of q_w and in that case it is one. So the derivative of the entire sum is always one regardless of w.</p>\n\n<p>The first part is a little bit more complex. When you take the derivative of  (Q_n -H(n \\ge v ))^2 you have to use the chain rule. First you work the ^2 which gives you 2 (Q_n -H(n \\ge v )) but then you have to take the derivative of Q_n with respect to q_w. </p>\n\n<p>Q_n is just the sum of q_j up to n. So if n is \\ge w then q_w will appear in this sum and the derivative will be  one. But if n is \\lt w then q_w will be absent and the derivative will be zero. Therefore the derivative of Q_n is just H(n \\ge w)</p>\n\n<p>HTH, Udi</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101939,
      "author_name": "nkhuyu",
      "author_url": "",
      "post_date": "12/18/2015 00:23:39",
      "content": "<p>That makes sense! Thank you very much!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101997,
      "author_name": "udibr1",
      "author_url": "",
      "post_date": "12/18/2015 13:02:54",
      "content": "<p>I guess javascript is disabled on your browser. You can read the document in text form at <a href=\"https://raw.githubusercontent.com/udibr/DSB2/master/optimal.ipynb\">https://raw.githubusercontent.com/udibr/DSB2/master/optimal.ipynb</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102001,
      "author_name": "hakimhakim",
      "author_url": "",
      "post_date": "12/18/2015 14:12:28",
      "content": "<p>[quote=udibr;101765]</p>\n\n<p>short math on why 0.042738 is the best score you can get without looking at the images\n<a href=\"https://github.com/udibr/DSB2/blob/master/optimal.ipynb\">link</a></p>\n\n<p>[/quote]</p>\n\n<p>Thank you udibr.\nI agree with you that the best score you can get without looking at the images is by using the cdf.\nBut the true cdf is not known and you use an empirical estimation in your code. So in my opinion and mathematically speaking, 0.042738 is not the best score because others estimations of the cdf exist. Anyway the best score is most likely very close to 0.042738.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102015,
      "author_name": "udibr1",
      "author_url": "",
      "post_date": "12/18/2015 16:49:37",
      "content": "<p>I agree. You can also overfit the validation set and reach even lower score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102057,
      "author_name": "jma127",
      "author_url": "",
      "post_date": "12/18/2015 23:47:22",
      "content": "<p>Hmm, assume the true distribution is a uniform distribution -- 1/600 for each volume bucket. Then, wouldn't the optimal guess for the CDF be 0 for 0 &lt; x &lt; 300, and 1 for x &gt;= 300, to get an expected score of 0.25 per sample? If you were to guess the true linear CDF, it seems you would get 0.33 per sample.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102065,
      "author_name": "udibr1",
      "author_url": "",
      "post_date": "12/19/2015 01:44:44",
      "content": "<p>for a step CDF you will get 0.5*(1/600)^2 + 0.5*(1-1/600)^2 = 0.499</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102066,
      "author_name": "jma127",
      "author_url": "",
      "post_date": "12/19/2015 01:46:24",
      "content": "<p>Oh derp, I didn't see the square term. My bad!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103954,
      "author_name": "udayabhanu",
      "author_url": "",
      "post_date": "01/08/2016 06:49:05",
      "content": "<p>Udibr,  shouldn't this be &gt;= instead of =? </p>\n\n<p>\\sum_{n=w} Q_n</p>\n\n<p>And, please explain to me how did you get 0.499 for step CDF, what is x there? Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104081,
      "author_name": "udibr1",
      "author_url": "",
      "post_date": "01/09/2016 00:13:54",
      "content": "<p>I didn't write the upward limit of the sum, M. The sum starts from n=w and runs upwards until M (so &gt;= would have been another way of saying it)\nI'm not sure I understand the second part of your question</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104289,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "01/11/2016 12:55:59",
      "content": "<p>Well, I'm already at 0.036361 without doing any image processing or machine learning. It is a gimmick, but it works well for testing my software infrastructure. I will improve it a little bit and make a post about it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "101765": "short math on why 0.042738 is the best score you can get without looking at the images\r\n[link][1]\r\n \r\n\r\n\r\n  [1]: https://github.com/udibr/DSB2/blob/master/optimal.ipynb",
    "101931": "Thank you!\r\n\r\nCan you show more details how do you take the derivative to get the following equation?\r\n\r\n${ \\partial C \\over \\partial q_w } = E_{p_v} \\left( \\sum_n 2 (Q_n -H(n \\ge v )) H(n \\ge w)\\right) + \\lambda",
    "101932": "Each q_j is a different variable and you have 600 of them. So you are actually making 600 different derivatives.  The part with the \\lambda in front of it is easy. In the sum each q_j appears once and when you decide to take a derivative according to just one of the q (for example q_w) then the derivative of each is part of the sum is zero unless you happen to take the derivative of q_w and in that case it is one. So the derivative of the entire sum is always one regardless of w.\r\n\r\nThe first part is a little bit more complex. When you take the derivative of  (Q_n -H(n \\ge v ))^2 you have to use the chain rule. First you work the ^2 which gives you 2 (Q_n -H(n \\ge v )) but then you have to take the derivative of Q_n with respect to q_w. \r\n\r\nQ_n is just the sum of q_j up to n. So if n is \\ge w then q_w will appear in this sum and the derivative will be  one. But if n is \\lt w then q_w will be absent and the derivative will be zero. Therefore the derivative of Q_n is just H(n \\ge w)\r\n\r\nHTH, Udi",
    "101939": "That makes sense! Thank you very much!",
    "101997": "I guess javascript is disabled on your browser. You can read the document in text form at https://raw.githubusercontent.com/udibr/DSB2/master/optimal.ipynb",
    "102001": "[quote=udibr;101765]\r\n\r\nshort math on why 0.042738 is the best score you can get without looking at the images\r\n[link][1]\r\n \r\n\r\n\r\n  [1]: https://github.com/udibr/DSB2/blob/master/optimal.ipynb\r\n\r\n[/quote]\r\n\r\nThank you udibr.\r\nI agree with you that the best score you can get without looking at the images is by using the cdf.\r\nBut the true cdf is not known and you use an empirical estimation in your code. So in my opinion and mathematically speaking, 0.042738 is not the best score because others estimations of the cdf exist. Anyway the best score is most likely very close to 0.042738.",
    "102015": "I agree. You can also overfit the validation set and reach even lower score.",
    "102057": "Hmm, assume the true distribution is a uniform distribution -- 1/600 for each volume bucket. Then, wouldn't the optimal guess for the CDF be 0 for 0 < x < 300, and 1 for x >= 300, to get an expected score of 0.25 per sample? If you were to guess the true linear CDF, it seems you would get 0.33 per sample.",
    "102065": "for a step CDF you will get 0.5*(1/600)^2 + 0.5*(1-1/600)^2 = 0.499",
    "102066": "Oh derp, I didn't see the square term. My bad!",
    "103954": "Udibr,  shouldn't this be >= instead of =? \r\n\r\n\\\\sum_{n=w} Q_n\r\n\r\nAnd, please explain to me how did you get 0.499 for step CDF, what is x there? Thanks.",
    "104081": "I didn't write the upward limit of the sum, M. The sum starts from n=w and runs upwards until M (so >= would have been another way of saying it)\r\nI'm not sure I understand the second part of your question",
    "104289": "Well, I'm already at 0.036361 without doing any image processing or machine learning. It is a gimmick, but it works well for testing my software infrastructure. I will improve it a little bit and make a post about it."
  },
  "source": "meta"
}