{
  "id": 2644,
  "title": "Multi Class Log Loss Function",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2644",
  "author_name": "",
  "post_date": "2012-09-12T01:31:59.257Z",
  "votes": 3,
  "comment_count": 5,
  "views": 7228,
  "content": "<p>Is the following Python code an accurate representation of how submissions are evaluated? I've played with this to help me evaluate my modelling, but wanted to make sure I understood how the evaluator worked. I believe I'll need to add a PostId to the prediction\r\n data when I submit, but have not included that for simplicity's sake in this example code.&nbsp;</p>\r\n<pre>from __future__ import division\r\n\r\nimport csv\r\nimport os\r\nimport scipy as sp\r\n\r\ndef llfun(act, pred):\r\n    epsilon = 1e-15\r\n    pred = sp.maximum(epsilon, pred)\r\n    pred = sp.minimum(1-epsilon, pred)\r\n    ll = sum(act*sp.log(pred) &#43; sp.subtract(1,act)*sp.log(sp.subtract(1,pred)))\r\n    ll = ll * -1.0/len(act)\r\n    return ll\r\n\r\ndef main():\r\n    pred = [\r\n        [0.05,0.05,0.05,0.8,0.05],\r\n        [0.73,0.05,0.01,0.20,0.02],\r\n        [0.02,0.03,0.01,0.75,0.19],\r\n        [0.01,0.02,0.83,0.12,0.02]\r\n        ]\r\n    act = [\r\n           [0,0,0,1,0],\r\n           [1,0,0,0,0],\r\n           [0,0,0,1,0],\r\n           [0,0,1,0,0]\r\n           ]\r\n\r\n    scores = []\r\n    for index in range(0, len(pred)):\r\n        result = llfun(act[index], pred[index])\r\n        scores.append(result)\r\n\r\n    print(sum(scores) / len(scores)) # 0.0985725708595\r\n\r\nif __name__ == '__main__':\r\n    main()\r\n</pre>",
  "messages": [
    {
      "id": "14247",
      "postDate": "09/12/2012 01:31:59",
      "content": "<p>Is the following Python code an accurate representation of how submissions are evaluated? I've played with this to help me evaluate my modelling, but wanted to make sure I understood how the evaluator worked. I believe I'll need to add a PostId to the prediction\r\n data when I submit, but have not included that for simplicity's sake in this example code.&nbsp;</p>\r\n<pre>from __future__ import division\r\n\r\nimport csv\r\nimport os\r\nimport scipy as sp\r\n\r\ndef llfun(act, pred):\r\n    epsilon = 1e-15\r\n    pred = sp.maximum(epsilon, pred)\r\n    pred = sp.minimum(1-epsilon, pred)\r\n    ll = sum(act*sp.log(pred) &#43; sp.subtract(1,act)*sp.log(sp.subtract(1,pred)))\r\n    ll = ll * -1.0/len(act)\r\n    return ll\r\n\r\ndef main():\r\n    pred = [\r\n        [0.05,0.05,0.05,0.8,0.05],\r\n        [0.73,0.05,0.01,0.20,0.02],\r\n        [0.02,0.03,0.01,0.75,0.19],\r\n        [0.01,0.02,0.83,0.12,0.02]\r\n        ]\r\n    act = [\r\n           [0,0,0,1,0],\r\n           [1,0,0,0,0],\r\n           [0,0,0,1,0],\r\n           [0,0,1,0,0]\r\n           ]\r\n\r\n    scores = []\r\n    for index in range(0, len(pred)):\r\n        result = llfun(act[index], pred[index])\r\n        scores.append(result)\r\n\r\n    print(sum(scores) / len(scores)) # 0.0985725708595\r\n\r\nif __name__ == '__main__':\r\n    main()\r\n</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14248",
      "postDate": "09/12/2012 02:36:25",
      "content": "<pre><code>ll = sum(act*sp.log(pred) &#43; sp.subtract(1,act)*sp.log(sp.subtract(1,pred)))\n</code></pre>\r\n<p>That's not right. Since the prediction is a normalized multinomial distribution, you just take log(pred[label]), and ignore the other predictions not covered by the label (their impact on the score is via the normalization). If your prediction is not actually\r\n normalized, you need to normalize it (after clamping to 1e-15).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14250",
      "postDate": "09/12/2012 07:34:58",
      "content": "<p>Here's the function I use:</p>\r\n<pre>import numpy as np\r\n\r\ndef multiclass_log_loss(y_true, y_pred, eps=1e-15):\r\n    &quot;&quot;&quot;Multi class version of Logarithmic Loss metric.\r\n    https://www.kaggle.com/wiki/MultiClassLogLoss\r\n\r\n    idea from this post:\r\n    http://www.kaggle.com/c/emc-data-science/forums/t/2149/is-anyone-noticing-difference-betwen-validation-and-leaderboard-error/12209#post12209\r\n\r\n    Parameters\r\n    ----------\r\n    y_true : array, shape = [n_samples]\r\n    y_pred : array, shape = [n_samples, n_classes]\r\n\r\n    Returns\r\n    -------\r\n    loss : float\r\n    &quot;&quot;&quot;\r\n    predictions = np.clip(y_pred, eps, 1 - eps)\r\n\r\n    # normalize row sums to 1\r\n    predictions /= predictions.sum(axis=1)[:, np.newaxis]\r\n\r\n    actual = np.zeros(y_pred.shape)\r\n    rows = actual.shape[0]\r\n    actual[np.arange(rows), y_true.astype(int)] = 1\r\n    vsota = np.sum(actual * np.log(predictions))\r\n    return -1.0 / rows * vsota\r\n</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14446",
      "postDate": "09/18/2012 00:24:58",
      "content": "<p>[quote=ephes;14250]</p>\r\n<p>Here's the function I use:</p>\r\n<pre>import numpy as np\r\n\r\ndef multiclass_log_loss(y_true, y_pred, eps=1e-15):\r\n    &quot;&quot;&quot;Multi class version of Logarithmic Loss metric.\r\n    https://www.kaggle.com/wiki/MultiClassLogLoss\r\n\r\n    idea from this post:\r\n    http://www.kaggle.com/c/emc-data-science/forums/t/2149/is-anyone-noticing-difference-betwen-validation-and-leaderboard-error/12209#post12209\r\n\r\n    Parameters\r\n    ----------\r\n    y_true : array, shape = [n_samples]\r\n    y_pred : array, shape = [n_samples, n_classes]\r\n\r\n    Returns\r\n    -------\r\n    loss : float\r\n    &quot;&quot;&quot;\r\n    predictions = np.clip(y_pred, eps, 1 - eps)\r\n\r\n    # normalize row sums to 1\r\n    predictions /= predictions.sum(axis=1)[:, np.newaxis]\r\n\r\n    actual = np.zeros(y_pred.shape)\r\n    rows = actual.shape[0]\r\n    actual[np.arange(rows), y_true.astype(int)] = 1\r\n    vsota = np.sum(actual * np.log(predictions))\r\n    return -1.0 / rows * vsota\r\n</pre>\r\n<p>[/quote]</p>\r\n<p>What type of objects are the inputs? &nbsp;</p>\r\n<p>&nbsp; &nbsp; &nbsp; y_true : array, shape = [n_samples]</p>\r\n<p>&nbsp; &nbsp; &nbsp; y_pred : array, shape = [n_samples, n_classes]</p>\r\n<p>&nbsp;</p>\r\n<p>I'm using a simple list of list and isn't working properly =(<br>\r\n<br>\r\nthanks for any help&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14531",
      "postDate": "09/19/2012 04:52:50",
      "content": "<p>The function assumes that two numpy ndarrays are supplied.</p>\r\n<p>The first is a 1-d array, where each element is the goldstandard class ID of the instance.</p>\r\n<p>The second is a 2-d array, where each element is the predicted distribution over the classes.</p>\r\n<p>Here are some example uses:</p>\r\n<pre>&gt;&gt;&gt; import numpy as np<br>&gt;&gt;&gt; multiclass_log_loss(np.array([0,1,2]),np.array([[1,0,0],[0,1,0],[0,0,1]]))<br>2.1094237467877998e-15<br>&gt;&gt;&gt; multiclass_log_loss(np.array([0,1,2]),np.array([[1,1,1],[0,1,0],[0,0,1]]))<br>0.36620409622270467</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "145265",
      "postDate": "11/17/2016 15:54:07",
      "content": "<p>The other option would be to use the built in log_loss function in scikit learn.  I've checked the results of that out of the box function against the ones posted above and it matches exactly.  just use</p>\n\n<pre><code>&gt;&gt;&gt; from sklearn.metrics import log_loss\n&gt;&gt;&gt; log_loss(np.array([0,1,2]),np.array([[1,0,0],[0,1,0],[0,0,1]]))\n2.1094237467877998e-15\n&gt;&gt;&gt; log_loss(np.array([0,1,2]),np.array([[1,1,1],[0,1,0],[0,0,1]]))\n0.36620409622270467\n</code></pre>\n\n<p>One of the advantages of the sklearn package is that you can leave your actual values coded as characters or strings.  It assumes that the probabilities in the prediction array are ordered alphabetically.  For example...</p>\n\n<pre><code>&gt;&gt;&gt; from sklearn.metrics import log_loss\n&gt;&gt;&gt; log_loss(np.array(['a','b','c']),np.array([[1,0,0],[0,1,0],[0,0,1]]))\n2.1094237467877998e-15\n&gt;&gt;&gt; log_loss(np.array(['a','b','c']),np.array([[1,1,1],[0,1,0],[0,0,1]]))\n0.36620409622270467\n&gt;&gt;&gt; log_loss(np.array(['cat','ant','boy']),np.array([[1,1,1],[0,1,0],[0,0,1]]))\n23.39205502616316\n</code></pre>",
      "rawMarkdown": "The other option would be to use the built in log_loss function in scikit learn.  I've checked the results of that out of the box function against the ones posted above and it matches exactly.  just use\r\n\r\n    >>> from sklearn.metrics import log_loss\r\n    >>> log_loss(np.array([0,1,2]),np.array([[1,0,0],[0,1,0],[0,0,1]]))\r\n    2.1094237467877998e-15\r\n    >>> log_loss(np.array([0,1,2]),np.array([[1,1,1],[0,1,0],[0,0,1]]))\r\n    0.36620409622270467\r\nOne of the advantages of the sklearn package is that you can leave your actual values coded as characters or strings.  It assumes that the probabilities in the prediction array are ordered alphabetically.  For example...\r\n\r\n    >>> from sklearn.metrics import log_loss\r\n    >>> log_loss(np.array(['a','b','c']),np.array([[1,0,0],[0,1,0],[0,0,1]]))\r\n    2.1094237467877998e-15\r\n    >>> log_loss(np.array(['a','b','c']),np.array([[1,1,1],[0,1,0],[0,0,1]]))\r\n    0.36620409622270467\r\n    >>> log_loss(np.array(['cat','ant','boy']),np.array([[1,1,1],[0,1,0],[0,0,1]]))\r\n    23.39205502616316",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 14248,
      "author_name": "andysloane",
      "author_url": "",
      "post_date": "09/12/2012 02:36:25",
      "content": "<pre><code>ll = sum(act*sp.log(pred) &#43; sp.subtract(1,act)*sp.log(sp.subtract(1,pred)))\n</code></pre>\r\n<p>That's not right. Since the prediction is a normalized multinomial distribution, you just take log(pred[label]), and ignore the other predictions not covered by the label (their impact on the score is via the normalization). If your prediction is not actually\r\n normalized, you need to normalize it (after clamping to 1e-15).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 14250,
      "author_name": "ephesus",
      "author_url": "",
      "post_date": "09/12/2012 07:34:58",
      "content": "<p>Here's the function I use:</p>\r\n<pre>import numpy as np\r\n\r\ndef multiclass_log_loss(y_true, y_pred, eps=1e-15):\r\n    &quot;&quot;&quot;Multi class version of Logarithmic Loss metric.\r\n    https://www.kaggle.com/wiki/MultiClassLogLoss\r\n\r\n    idea from this post:\r\n    http://www.kaggle.com/c/emc-data-science/forums/t/2149/is-anyone-noticing-difference-betwen-validation-and-leaderboard-error/12209#post12209\r\n\r\n    Parameters\r\n    ----------\r\n    y_true : array, shape = [n_samples]\r\n    y_pred : array, shape = [n_samples, n_classes]\r\n\r\n    Returns\r\n    -------\r\n    loss : float\r\n    &quot;&quot;&quot;\r\n    predictions = np.clip(y_pred, eps, 1 - eps)\r\n\r\n    # normalize row sums to 1\r\n    predictions /= predictions.sum(axis=1)[:, np.newaxis]\r\n\r\n    actual = np.zeros(y_pred.shape)\r\n    rows = actual.shape[0]\r\n    actual[np.arange(rows), y_true.astype(int)] = 1\r\n    vsota = np.sum(actual * np.log(predictions))\r\n    return -1.0 / rows * vsota\r\n</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 14446,
      "author_name": "alessandrosena",
      "author_url": "",
      "post_date": "09/18/2012 00:24:58",
      "content": "<p>[quote=ephes;14250]</p>\r\n<p>Here's the function I use:</p>\r\n<pre>import numpy as np\r\n\r\ndef multiclass_log_loss(y_true, y_pred, eps=1e-15):\r\n    &quot;&quot;&quot;Multi class version of Logarithmic Loss metric.\r\n    https://www.kaggle.com/wiki/MultiClassLogLoss\r\n\r\n    idea from this post:\r\n    http://www.kaggle.com/c/emc-data-science/forums/t/2149/is-anyone-noticing-difference-betwen-validation-and-leaderboard-error/12209#post12209\r\n\r\n    Parameters\r\n    ----------\r\n    y_true : array, shape = [n_samples]\r\n    y_pred : array, shape = [n_samples, n_classes]\r\n\r\n    Returns\r\n    -------\r\n    loss : float\r\n    &quot;&quot;&quot;\r\n    predictions = np.clip(y_pred, eps, 1 - eps)\r\n\r\n    # normalize row sums to 1\r\n    predictions /= predictions.sum(axis=1)[:, np.newaxis]\r\n\r\n    actual = np.zeros(y_pred.shape)\r\n    rows = actual.shape[0]\r\n    actual[np.arange(rows), y_true.astype(int)] = 1\r\n    vsota = np.sum(actual * np.log(predictions))\r\n    return -1.0 / rows * vsota\r\n</pre>\r\n<p>[/quote]</p>\r\n<p>What type of objects are the inputs? &nbsp;</p>\r\n<p>&nbsp; &nbsp; &nbsp; y_true : array, shape = [n_samples]</p>\r\n<p>&nbsp; &nbsp; &nbsp; y_pred : array, shape = [n_samples, n_classes]</p>\r\n<p>&nbsp;</p>\r\n<p>I'm using a simple list of list and isn't working properly =(<br>\r\n<br>\r\nthanks for any help&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 14531,
      "author_name": "marcolui",
      "author_url": "",
      "post_date": "09/19/2012 04:52:50",
      "content": "<p>The function assumes that two numpy ndarrays are supplied.</p>\r\n<p>The first is a 1-d array, where each element is the goldstandard class ID of the instance.</p>\r\n<p>The second is a 2-d array, where each element is the predicted distribution over the classes.</p>\r\n<p>Here are some example uses:</p>\r\n<pre>&gt;&gt;&gt; import numpy as np<br>&gt;&gt;&gt; multiclass_log_loss(np.array([0,1,2]),np.array([[1,0,0],[0,1,0],[0,0,1]]))<br>2.1094237467877998e-15<br>&gt;&gt;&gt; multiclass_log_loss(np.array([0,1,2]),np.array([[1,1,1],[0,1,0],[0,0,1]]))<br>0.36620409622270467</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145265,
      "author_name": "jedisom",
      "author_url": "",
      "post_date": "11/17/2016 15:54:07",
      "content": "<p>The other option would be to use the built in log_loss function in scikit learn.  I've checked the results of that out of the box function against the ones posted above and it matches exactly.  just use</p>\n\n<pre><code>&gt;&gt;&gt; from sklearn.metrics import log_loss\n&gt;&gt;&gt; log_loss(np.array([0,1,2]),np.array([[1,0,0],[0,1,0],[0,0,1]]))\n2.1094237467877998e-15\n&gt;&gt;&gt; log_loss(np.array([0,1,2]),np.array([[1,1,1],[0,1,0],[0,0,1]]))\n0.36620409622270467\n</code></pre>\n\n<p>One of the advantages of the sklearn package is that you can leave your actual values coded as characters or strings.  It assumes that the probabilities in the prediction array are ordered alphabetically.  For example...</p>\n\n<pre><code>&gt;&gt;&gt; from sklearn.metrics import log_loss\n&gt;&gt;&gt; log_loss(np.array(['a','b','c']),np.array([[1,0,0],[0,1,0],[0,0,1]]))\n2.1094237467877998e-15\n&gt;&gt;&gt; log_loss(np.array(['a','b','c']),np.array([[1,1,1],[0,1,0],[0,0,1]]))\n0.36620409622270467\n&gt;&gt;&gt; log_loss(np.array(['cat','ant','boy']),np.array([[1,1,1],[0,1,0],[0,0,1]]))\n23.39205502616316\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "14247": "",
    "14248": "",
    "14250": "",
    "14446": "",
    "14531": "",
    "145265": "The other option would be to use the built in log_loss function in scikit learn.  I've checked the results of that out of the box function against the ones posted above and it matches exactly.  just use\r\n\r\n    >>> from sklearn.metrics import log_loss\r\n    >>> log_loss(np.array([0,1,2]),np.array([[1,0,0],[0,1,0],[0,0,1]]))\r\n    2.1094237467877998e-15\r\n    >>> log_loss(np.array([0,1,2]),np.array([[1,1,1],[0,1,0],[0,0,1]]))\r\n    0.36620409622270467\r\nOne of the advantages of the sklearn package is that you can leave your actual values coded as characters or strings.  It assumes that the probabilities in the prediction array are ordered alphabetically.  For example...\r\n\r\n    >>> from sklearn.metrics import log_loss\r\n    >>> log_loss(np.array(['a','b','c']),np.array([[1,0,0],[0,1,0],[0,0,1]]))\r\n    2.1094237467877998e-15\r\n    >>> log_loss(np.array(['a','b','c']),np.array([[1,1,1],[0,1,0],[0,0,1]]))\r\n    0.36620409622270467\r\n    >>> log_loss(np.array(['cat','ant','boy']),np.array([[1,1,1],[0,1,0],[0,0,1]]))\r\n    23.39205502616316"
  },
  "source": "meta"
}