{
  "id": 33072,
  "title": "Log Loss calculation",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/33072",
  "author_name": "",
  "post_date": "2017-05-16T11:00:41.191477400Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have a question with how the log loss is calculated. Please excuse me if this is a simple question but i'm just trying to understand :)</p>\n\n<p>I read here:\n<a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening#evaluation\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening#evaluation</a></p>\n\n<p>that the ground truth is one of the three classes and not split across classes \" 1 if observation (i) belongs to class (j) and 0 otherwise\". In that case the equation given there ignores the classes that are zero. (i.e. if yi = [1,0,0], the 0's, or really small number, eliminate the wrong guess) So the log of the prediction for the correct class is all that's considered. Hence if I were to put in a submission of [0.33,0.33,0.33] for all 512 images I would expect an Average score of  0.47700 (i.e. (1)*log(0.333) = 0.477).</p>\n\n<p>however, this formula is actually different to link to multi-class log loss (which is almost the same as cross-entropy in this case) <a href=\"https://www.kaggle.com/wiki/MultiClassLogLoss\">https://www.kaggle.com/wiki/MultiClassLogLoss</a>\nbut with this equation putting in [0.33,0.33,0.33] for all 512 images would give 0.6365</p>\n\n<p>from sklearn.metrics import log_loss\nprediction = [0.333333,0.33333,0.333333]\ngroundtruth = [1,0,0]</p>\n\n<p>log_loss(groundtruth, prediction, eps=1e-15)\nOut[1]: 0.63651266829918784</p>\n\n<p>The main reason why i want to know how this score is actually calculated is that we are using it as the metric to assess whether a change we make to our models is a good change or not.</p>\n\n<p>Please help if i'm missing something really simple just so i can understand</p>",
  "messages": [
    {
      "id": "182921",
      "postDate": "05/16/2017 11:00:41",
      "content": "<p>I have a question with how the log loss is calculated. Please excuse me if this is a simple question but i'm just trying to understand :)</p>\n\n<p>I read here:\n<a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening#evaluation\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening#evaluation</a></p>\n\n<p>that the ground truth is one of the three classes and not split across classes \" 1 if observation (i) belongs to class (j) and 0 otherwise\". In that case the equation given there ignores the classes that are zero. (i.e. if yi = [1,0,0], the 0's, or really small number, eliminate the wrong guess) So the log of the prediction for the correct class is all that's considered. Hence if I were to put in a submission of [0.33,0.33,0.33] for all 512 images I would expect an Average score of  0.47700 (i.e. (1)*log(0.333) = 0.477).</p>\n\n<p>however, this formula is actually different to link to multi-class log loss (which is almost the same as cross-entropy in this case) <a href=\"https://www.kaggle.com/wiki/MultiClassLogLoss\">https://www.kaggle.com/wiki/MultiClassLogLoss</a>\nbut with this equation putting in [0.33,0.33,0.33] for all 512 images would give 0.6365</p>\n\n<p>from sklearn.metrics import log_loss\nprediction = [0.333333,0.33333,0.333333]\ngroundtruth = [1,0,0]</p>\n\n<p>log_loss(groundtruth, prediction, eps=1e-15)\nOut[1]: 0.63651266829918784</p>\n\n<p>The main reason why i want to know how this score is actually calculated is that we are using it as the metric to assess whether a change we make to our models is a good change or not.</p>\n\n<p>Please help if i'm missing something really simple just so i can understand</p>",
      "rawMarkdown": "I have a question with how the log loss is calculated. Please excuse me if this is a simple question but i'm just trying to understand :)\n\nI read here:\nhttps://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening#evaluation\n\nthat the ground truth is one of the three classes and not split across classes \" 1 if observation (i) belongs to class (j) and 0 otherwise\". In that case the equation given there ignores the classes that are zero. (i.e. if yi = [1,0,0], the 0's, or really small number, eliminate the wrong guess) So the log of the prediction for the correct class is all that's considered. Hence if I were to put in a submission of [0.33,0.33,0.33] for all 512 images I would expect an Average score of  0.47700 (i.e. (1)*log(0.333) = 0.477).\n\nhowever, this formula is actually different to link to multi-class log loss (which is almost the same as cross-entropy in this case) https://www.kaggle.com/wiki/MultiClassLogLoss\nbut with this equation putting in [0.33,0.33,0.33] for all 512 images would give 0.6365\n\nfrom sklearn.metrics import log_loss\nprediction = [0.333333,0.33333,0.333333]\ngroundtruth = [1,0,0]\n\nlog_loss(groundtruth, prediction, eps=1e-15)\nOut[1]: 0.63651266829918784\n\nThe main reason why i want to know how this score is actually calculated is that we are using it as the metric to assess whether a change we make to our models is a good change or not.\n\nPlease help if i'm missing something really simple just so i can understand",
      "votes": null
    },
    {
      "id": "182923",
      "postDate": "05/16/2017 11:16:54",
      "content": "<p>I think it's using natural logs rather than to the base 10? So a submission with 0.333 for each type would yield a score of -LN(0.333) = 1.1087, and that the formula being used ignores the classes that are zero and takes an average of -LN(predicted prob) for the classes that are 1.</p>",
      "rawMarkdown": "I think it's using natural logs rather than to the base 10? So a submission with 0.333 for each type would yield a score of -LN(0.333) = 1.1087, and that the formula being used ignores the classes that are zero and takes an average of -LN(predicted prob) for the classes that are 1.",
      "votes": null
    },
    {
      "id": "182925",
      "postDate": "05/16/2017 11:27:36",
      "content": "<p>... yup... like i thought it was something silly that i'd missed :) \nmany thanks fergusoci :)</p>",
      "rawMarkdown": "... yup... like i thought it was something silly that i'd missed :) \nmany thanks fergusoci :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 182923,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "05/16/2017 11:16:54",
      "content": "<p>I think it's using natural logs rather than to the base 10? So a submission with 0.333 for each type would yield a score of -LN(0.333) = 1.1087, and that the formula being used ignores the classes that are zero and takes an average of -LN(predicted prob) for the classes that are 1.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 182925,
      "author_name": "monodavids",
      "author_url": "",
      "post_date": "05/16/2017 11:27:36",
      "content": "<p>... yup... like i thought it was something silly that i'd missed :) \nmany thanks fergusoci :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "182921": "I have a question with how the log loss is calculated. Please excuse me if this is a simple question but i'm just trying to understand :)\n\nI read here:\nhttps://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening#evaluation\n\nthat the ground truth is one of the three classes and not split across classes \" 1 if observation (i) belongs to class (j) and 0 otherwise\". In that case the equation given there ignores the classes that are zero. (i.e. if yi = [1,0,0], the 0's, or really small number, eliminate the wrong guess) So the log of the prediction for the correct class is all that's considered. Hence if I were to put in a submission of [0.33,0.33,0.33] for all 512 images I would expect an Average score of  0.47700 (i.e. (1)*log(0.333) = 0.477).\n\nhowever, this formula is actually different to link to multi-class log loss (which is almost the same as cross-entropy in this case) https://www.kaggle.com/wiki/MultiClassLogLoss\nbut with this equation putting in [0.33,0.33,0.33] for all 512 images would give 0.6365\n\nfrom sklearn.metrics import log_loss\nprediction = [0.333333,0.33333,0.333333]\ngroundtruth = [1,0,0]\n\nlog_loss(groundtruth, prediction, eps=1e-15)\nOut[1]: 0.63651266829918784\n\nThe main reason why i want to know how this score is actually calculated is that we are using it as the metric to assess whether a change we make to our models is a good change or not.\n\nPlease help if i'm missing something really simple just so i can understand",
    "182923": "I think it's using natural logs rather than to the base 10? So a submission with 0.333 for each type would yield a score of -LN(0.333) = 1.1087, and that the formula being used ignores the classes that are zero and takes an average of -LN(predicted prob) for the classes that are 1.",
    "182925": "... yup... like i thought it was something silly that i'd missed :) \nmany thanks fergusoci :)"
  },
  "source": "meta"
}