{
  "id": 5219,
  "title": "evaluation formula",
  "url": "/competitions/belkin-energy-disaggregation-competition/discussion/5219",
  "author_name": "",
  "post_date": "2013-07-28T17:33:30.460Z",
  "votes": null,
  "comment_count": 5,
  "views": 2039,
  "content": "<p>Hi,</p>\n<p>I wonder if the evaluation formula (Hamming Loss) is correct.</p>\n<p>A score of 0.07809 would mean that approximately 90 % of the</p>\n<p>'All Appliances Always Off Benchmark' predictions are false.</p>\n<p>A 'All Appliances Always On Benchmark' would score 0.00786.</p>\n<p>What am I missing? Should the Predicted Values be the Appliance (instead</p>\n<p>of 1 for ON and 0 for OFF)? Or is the evaluation formula wrong (XNOR instead XOR)?</p>",
  "messages": [
    {
      "id": "27732",
      "postDate": "07/28/2013 17:33:30",
      "content": "<p>Hi,</p>\n<p>I wonder if the evaluation formula (Hamming Loss) is correct.</p>\n<p>A score of 0.07809 would mean that approximately 90 % of the</p>\n<p>'All Appliances Always Off Benchmark' predictions are false.</p>\n<p>A 'All Appliances Always On Benchmark' would score 0.00786.</p>\n<p>What am I missing? Should the Predicted Values be the Appliance (instead</p>\n<p>of 1 for ON and 0 for OFF)? Or is the evaluation formula wrong (XNOR instead XOR)?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28447",
      "postDate": "08/08/2013 18:25:54",
      "content": "<p>The metric here penalizes misclassifications&nbsp;too much. For example, there 4 appliances in which only appliance 1 is on, that is 1,0,0,0 (ground true)</p>\n<p>All-off one gives 0,0,0,0. The loss is 1.</p>\n<p>A wrong estimation, e.g., 0,0,0,1. The loss is 2.</p>\n<p>This tell us that only classifiers whose accuracy is above 67%&nbsp;can beat the all-off benchmark.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28448",
      "postDate": "08/08/2013 18:35:08",
      "content": "<p>[quote=Guocong Song;28447]</p>\n<p>The metric here penalizes misclassifications&nbsp;too much. For example, there 4 appliances in which only appliance 1 is on, that is 1,0,0,0 (ground true)</p>\n<p>All-off one gives 0,0,0,0. The loss is 1.</p>\n<p>A wrong estimation, e.g., 0,0,0,1. The loss is 2.</p>\n<p>This tell us that only classifiers whose accuracy is above 67%&nbsp;can beat the all-off benchmark.</p>\n<p>[/quote]</p>\n<p>If there are 4 appliances and one time point and the ground truth is 1,0,0,0, the Hamming Loss of 0,0,0,0&nbsp;is:</p>\n<p>1/1 * (1+0+0+0)/4 = 1/4</p>\n<p>The Hamming Loss of&nbsp;0,0,0,1 is&nbsp;</p>\n<p>1/1 * (1+0+0+1)/4 = 1/2</p>\n<p>The first case made one mistake and the second case made two mistakes. Why would you expect anything but a doubling of the Hamming Loss?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28452",
      "postDate": "08/08/2013 19:28:27",
      "content": "<p>[quote=jessica bombaz;27732]</p>\n<p>Hi,</p>\n<p>I wonder if the evaluation formula (Hamming Loss) is correct.</p>\n<p>A score of 0.07809 would mean that approximately 90 % of the</p>\n<p>'All Appliances Always Off Benchmark' predictions are false.</p>\n<p>A 'All Appliances Always On Benchmark' would score 0.00786.</p>\n<p>What am I missing? Should the Predicted Values be the Appliance (instead</p>\n<p>of 1 for ON and 0 for OFF)? Or is the evaluation formula wrong (XNOR instead XOR)?</p>\n<p>[/quote]</p>\n<p>Jessica, this is a loss function, so a score closer to 0 is better.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "30952",
      "postDate": "09/16/2013 01:08:54",
      "content": "<p>By having each row correspond to a given appliance, then each row for a given house counts the same, right? (this is Noam's point 3 <a href=\"http://www.kaggle.com/c/belkin-energy-disaggregation-competition/forums/t/5579/some-assumptions\" target=\"_blank\">here</a>)</p>\n<p>If all appliances off gives a score ~8%, then that would mean ~8% of rows are incorrect. But I thought we were told that only one appliance would be on at one time, which with 36-38 appliances per home would conflict with this. (Since then even if one appliance is always running, only 1/38 - 1/36 of rows would be &quot;on&quot;?)</p>\n<p>Or did that rule only hold in the training data, and the test data has multiple appliances operating at once?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "30961",
      "postDate": "09/16/2013 09:05:26",
      "content": "<p>You are right &quot; that rule only holds in the training data, and the test data has multiple appliances operating at once&quot;. &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 28447,
      "author_name": "songgc",
      "author_url": "",
      "post_date": "08/08/2013 18:25:54",
      "content": "<p>The metric here penalizes misclassifications&nbsp;too much. For example, there 4 appliances in which only appliance 1 is on, that is 1,0,0,0 (ground true)</p>\n<p>All-off one gives 0,0,0,0. The loss is 1.</p>\n<p>A wrong estimation, e.g., 0,0,0,1. The loss is 2.</p>\n<p>This tell us that only classifiers whose accuracy is above 67%&nbsp;can beat the all-off benchmark.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28448,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "08/08/2013 18:35:08",
      "content": "<p>[quote=Guocong Song;28447]</p>\n<p>The metric here penalizes misclassifications&nbsp;too much. For example, there 4 appliances in which only appliance 1 is on, that is 1,0,0,0 (ground true)</p>\n<p>All-off one gives 0,0,0,0. The loss is 1.</p>\n<p>A wrong estimation, e.g., 0,0,0,1. The loss is 2.</p>\n<p>This tell us that only classifiers whose accuracy is above 67%&nbsp;can beat the all-off benchmark.</p>\n<p>[/quote]</p>\n<p>If there are 4 appliances and one time point and the ground truth is 1,0,0,0, the Hamming Loss of 0,0,0,0&nbsp;is:</p>\n<p>1/1 * (1+0+0+0)/4 = 1/4</p>\n<p>The Hamming Loss of&nbsp;0,0,0,1 is&nbsp;</p>\n<p>1/1 * (1+0+0+1)/4 = 1/2</p>\n<p>The first case made one mistake and the second case made two mistakes. Why would you expect anything but a doubling of the Hamming Loss?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28452,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "08/08/2013 19:28:27",
      "content": "<p>[quote=jessica bombaz;27732]</p>\n<p>Hi,</p>\n<p>I wonder if the evaluation formula (Hamming Loss) is correct.</p>\n<p>A score of 0.07809 would mean that approximately 90 % of the</p>\n<p>'All Appliances Always Off Benchmark' predictions are false.</p>\n<p>A 'All Appliances Always On Benchmark' would score 0.00786.</p>\n<p>What am I missing? Should the Predicted Values be the Appliance (instead</p>\n<p>of 1 for ON and 0 for OFF)? Or is the evaluation formula wrong (XNOR instead XOR)?</p>\n<p>[/quote]</p>\n<p>Jessica, this is a loss function, so a score closer to 0 is better.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 30952,
      "author_name": "rosnfeld",
      "author_url": "",
      "post_date": "09/16/2013 01:08:54",
      "content": "<p>By having each row correspond to a given appliance, then each row for a given house counts the same, right? (this is Noam's point 3 <a href=\"http://www.kaggle.com/c/belkin-energy-disaggregation-competition/forums/t/5579/some-assumptions\" target=\"_blank\">here</a>)</p>\n<p>If all appliances off gives a score ~8%, then that would mean ~8% of rows are incorrect. But I thought we were told that only one appliance would be on at one time, which with 36-38 appliances per home would conflict with this. (Since then even if one appliance is always running, only 1/38 - 1/36 of rows would be &quot;on&quot;?)</p>\n<p>Or did that rule only hold in the training data, and the test data has multiple appliances operating at once?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 30961,
      "author_name": "noamtene",
      "author_url": "",
      "post_date": "09/16/2013 09:05:26",
      "content": "<p>You are right &quot; that rule only holds in the training data, and the test data has multiple appliances operating at once&quot;. &nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "27732": "",
    "28447": "",
    "28448": "",
    "28452": "",
    "30952": "",
    "30961": ""
  },
  "source": "meta"
}