{
  "id": 18015,
  "title": "Competition evaluation metric for Python ",
  "url": "/competitions/second-annual-data-science-bowl/discussion/18015",
  "author_name": "",
  "post_date": "2015-12-18T22:21:13.933Z",
  "votes": 4,
  "comment_count": 4,
  "views": 1540,
  "content": "<p>In case anyone is interested. Feel free to suggest improvements or catch anything we missed. Using this code for model validation. </p>\n\n<pre><code>#..written for python 2.7\nimport numpy as np \nimport pandas as pd\n\ndef CRPS_row(row):\n    &quot;&quot;&quot;\n    This function is purpose-built for Kaggle. With a couple tweaks it can \n    be generalized for other applications. \n\n    row should be a 601-element list where the first element is the true volme \n    and the subsequent elements are cumulatively summed probabilities. \n    &quot;&quot;&quot;\n    V_m = row[0]\n    p = np.array(row[1:])\n    v = np.array(range(len(n)))\n    h = v &gt;= V_m\n    sq_dists = (p - h)**2\n    return(np.sum(sq_dists)/len(sq_dists)) \n\ndef CRPS_mean(df): \n    &quot;&quot;&quot;\n    Function recieves pandas dataframe as an input with the first column being\n    a column of truths (V_m) and the subsequent 600 columns being cumulatively\n    summed probabilities\n    &quot;&quot;&quot;\n    crps_vec = df.apply(CRPS_row, axis = 1)\n    crps_sum = np.sum(crps_vec)\n    return(crps_sum/len(crps_vec))\n</code></pre>",
  "messages": [
    {
      "id": "102050",
      "postDate": "12/18/2015 22:21:13",
      "content": "<p>In case anyone is interested. Feel free to suggest improvements or catch anything we missed. Using this code for model validation. </p>\n\n<pre><code>#..written for python 2.7\nimport numpy as np \nimport pandas as pd\n\ndef CRPS_row(row):\n    &quot;&quot;&quot;\n    This function is purpose-built for Kaggle. With a couple tweaks it can \n    be generalized for other applications. \n\n    row should be a 601-element list where the first element is the true volme \n    and the subsequent elements are cumulatively summed probabilities. \n    &quot;&quot;&quot;\n    V_m = row[0]\n    p = np.array(row[1:])\n    v = np.array(range(len(n)))\n    h = v &gt;= V_m\n    sq_dists = (p - h)**2\n    return(np.sum(sq_dists)/len(sq_dists)) \n\ndef CRPS_mean(df): \n    &quot;&quot;&quot;\n    Function recieves pandas dataframe as an input with the first column being\n    a column of truths (V_m) and the subsequent 600 columns being cumulatively\n    summed probabilities\n    &quot;&quot;&quot;\n    crps_vec = df.apply(CRPS_row, axis = 1)\n    crps_sum = np.sum(crps_vec)\n    return(crps_sum/len(crps_vec))\n</code></pre>",
      "rawMarkdown": "In case anyone is interested. Feel free to suggest improvements or catch anything we missed. Using this code for model validation. \r\n\r\n    #..written for python 2.7\r\n    import numpy as np \r\n    import pandas as pd\r\n    \r\n    def CRPS_row(row):\r\n        \"\"\"\r\n        This function is purpose-built for Kaggle. With a couple tweaks it can \r\n        be generalized for other applications. \r\n        \r\n        row should be a 601-element list where the first element is the true volme \r\n        and the subsequent elements are cumulatively summed probabilities. \r\n        \"\"\"\r\n        V_m = row[0]\r\n        p = np.array(row[1:])\r\n        v = np.array(range(len(n)))\r\n        h = v >= V_m\r\n        sq_dists = (p - h)**2\r\n        return(np.sum(sq_dists)/len(sq_dists)) \r\n    \r\n    def CRPS_mean(df): \r\n        \"\"\"\r\n        Function recieves pandas dataframe as an input with the first column being\r\n        a column of truths (V_m) and the subsequent 600 columns being cumulatively\r\n        summed probabilities\r\n        \"\"\"\r\n        crps_vec = df.apply(CRPS_row, axis = 1)\r\n        crps_sum = np.sum(crps_vec)\r\n        return(crps_sum/len(crps_vec))",
      "votes": null
    },
    {
      "id": "104380",
      "postDate": "01/12/2016 10:14:32",
      "content": "<p>Hi Aaron,</p>\n\n<p>Can you run evaluation of all zeros CDFs against <strong>train.csv</strong> set? I'm getting 0.777719 with my implementation and I would like to verify that it's accurate.</p>\n\n<p>Thanks,\nPaul</p>",
      "rawMarkdown": "Hi Aaron,\r\n\r\nCan you run evaluation of all zeros CDFs against **train.csv** set? I'm getting 0.777719 with my implementation and I would like to verify that it's accurate.\r\n\r\nThanks,\r\nPaul",
      "votes": null
    },
    {
      "id": "105300",
      "postDate": "01/21/2016 21:40:00",
      "content": "<p>Thanks for the above implementation.  I used it to perform a sanity check on my CRPS function. Against all zero cdfs on <strong>train.csv</strong> I got a score of <code>0.801195</code> for both implementations.   Paul's score still has me skeptical though.  </p>\n\n<p>I had to make a couple edits to <code>CRPS_row</code> to get reasonable results on my machine (py 2.7):</p>\n\n<pre><code>V_m = row.iloc[0]\n...\nv = np.array(range(len(p)))\n</code></pre>",
      "rawMarkdown": "Thanks for the above implementation.  I used it to perform a sanity check on my CRPS function. Against all zero cdfs on __train.csv__ I got a score of `0.801195` for both implementations.   Paul's score still has me skeptical though.  \r\n\r\nI had to make a couple edits to `CRPS_row` to get reasonable results on my machine (py 2.7):\r\n\r\n    V_m = row.iloc[0]\r\n    ...\r\n    v = np.array(range(len(p)))",
      "votes": null
    },
    {
      "id": "105408",
      "postDate": "01/22/2016 22:05:04",
      "content": "<p>Thanks Blake, good updates. Paul, I haven't really been able to Kaggle for the past couple weeks--sorry to leave you hanging. Sounds like Blake sanity checked the all zeros set two different ways and came up with the same result. I'd be inclined to trust it, but I'll let you know if I see anything different!</p>",
      "rawMarkdown": "Thanks Blake, good updates. Paul, I haven't really been able to Kaggle for the past couple weeks--sorry to leave you hanging. Sounds like Blake sanity checked the all zeros set two different ways and came up with the same result. I'd be inclined to trust it, but I'll let you know if I see anything different!",
      "votes": null
    },
    {
      "id": "105413",
      "postDate": "01/22/2016 23:14:16",
      "content": "<p>@Blake @Aaron</p>\n\n<p>I will have another look at my implementation to find an error somewhere.</p>",
      "rawMarkdown": "Blake @Aaron\r\n\r\nI will have another look at my implementation to find an error somewhere.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 104380,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "01/12/2016 10:14:32",
      "content": "<p>Hi Aaron,</p>\n\n<p>Can you run evaluation of all zeros CDFs against <strong>train.csv</strong> set? I'm getting 0.777719 with my implementation and I would like to verify that it's accurate.</p>\n\n<p>Thanks,\nPaul</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105300,
      "author_name": "blakeboswell",
      "author_url": "",
      "post_date": "01/21/2016 21:40:00",
      "content": "<p>Thanks for the above implementation.  I used it to perform a sanity check on my CRPS function. Against all zero cdfs on <strong>train.csv</strong> I got a score of <code>0.801195</code> for both implementations.   Paul's score still has me skeptical though.  </p>\n\n<p>I had to make a couple edits to <code>CRPS_row</code> to get reasonable results on my machine (py 2.7):</p>\n\n<pre><code>V_m = row.iloc[0]\n...\nv = np.array(range(len(p)))\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105408,
      "author_name": "aaronpolhamus",
      "author_url": "",
      "post_date": "01/22/2016 22:05:04",
      "content": "<p>Thanks Blake, good updates. Paul, I haven't really been able to Kaggle for the past couple weeks--sorry to leave you hanging. Sounds like Blake sanity checked the all zeros set two different ways and came up with the same result. I'd be inclined to trust it, but I'll let you know if I see anything different!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105413,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "01/22/2016 23:14:16",
      "content": "<p>@Blake @Aaron</p>\n\n<p>I will have another look at my implementation to find an error somewhere.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "102050": "In case anyone is interested. Feel free to suggest improvements or catch anything we missed. Using this code for model validation. \r\n\r\n    #..written for python 2.7\r\n    import numpy as np \r\n    import pandas as pd\r\n    \r\n    def CRPS_row(row):\r\n        \"\"\"\r\n        This function is purpose-built for Kaggle. With a couple tweaks it can \r\n        be generalized for other applications. \r\n        \r\n        row should be a 601-element list where the first element is the true volme \r\n        and the subsequent elements are cumulatively summed probabilities. \r\n        \"\"\"\r\n        V_m = row[0]\r\n        p = np.array(row[1:])\r\n        v = np.array(range(len(n)))\r\n        h = v >= V_m\r\n        sq_dists = (p - h)**2\r\n        return(np.sum(sq_dists)/len(sq_dists)) \r\n    \r\n    def CRPS_mean(df): \r\n        \"\"\"\r\n        Function recieves pandas dataframe as an input with the first column being\r\n        a column of truths (V_m) and the subsequent 600 columns being cumulatively\r\n        summed probabilities\r\n        \"\"\"\r\n        crps_vec = df.apply(CRPS_row, axis = 1)\r\n        crps_sum = np.sum(crps_vec)\r\n        return(crps_sum/len(crps_vec))",
    "104380": "Hi Aaron,\r\n\r\nCan you run evaluation of all zeros CDFs against **train.csv** set? I'm getting 0.777719 with my implementation and I would like to verify that it's accurate.\r\n\r\nThanks,\r\nPaul",
    "105300": "Thanks for the above implementation.  I used it to perform a sanity check on my CRPS function. Against all zero cdfs on __train.csv__ I got a score of `0.801195` for both implementations.   Paul's score still has me skeptical though.  \r\n\r\nI had to make a couple edits to `CRPS_row` to get reasonable results on my machine (py 2.7):\r\n\r\n    V_m = row.iloc[0]\r\n    ...\r\n    v = np.array(range(len(p)))",
    "105408": "Thanks Blake, good updates. Paul, I haven't really been able to Kaggle for the past couple weeks--sorry to leave you hanging. Sounds like Blake sanity checked the all zeros set two different ways and came up with the same result. I'd be inclined to trust it, but I'll let you know if I see anything different!",
    "105413": "Blake @Aaron\r\n\r\nI will have another look at my implementation to find an error somewhere."
  },
  "source": "meta"
}