{
  "id": 5317,
  "title": "Question about the evaluation function",
  "url": "/competitions/belkin-energy-disaggregation-competition/discussion/5317",
  "author_name": "",
  "post_date": "2013-08-06T16:06:49.323Z",
  "votes": null,
  "comment_count": 10,
  "views": 2825,
  "content": "<p>I am trying to understand how to apply the Hamming Loss formula given in the Evaluation section to the scores being generated for the sample submission.</p>\n<p>The sample submission has 1762 minutes for H1 with each minute having an entry for all 38 appliances. &nbsp;From this I conclude that D=1762 and L=38 and that the contribution of each H1 line to the H1 score is 1/(1762*38).</p>\n<p>Since the houses all get the same weight the contribution of each H1 line to the score should be 1/(4*1762*38) = 0.000003733</p>\n<p>I was therefore rather surprised to find out that my score for a&nbsp;<span style=\"line-height: 1.4\">submission which had only one appliance turned on for 92 minutes on H1 was 0.00068 better than the all appliances off benchmark.</span></p>\n<p><span style=\"line-height: 1.4\">I expected the maximum improvement to be 92*0.000003733 = 0.00034 and my score improved by twice that much.</span></p>\n<p>What am I missing?</p>",
  "messages": [
    {
      "id": "28297",
      "postDate": "08/06/2013 16:06:49",
      "content": "<p>I am trying to understand how to apply the Hamming Loss formula given in the Evaluation section to the scores being generated for the sample submission.</p>\n<p>The sample submission has 1762 minutes for H1 with each minute having an entry for all 38 appliances. &nbsp;From this I conclude that D=1762 and L=38 and that the contribution of each H1 line to the H1 score is 1/(1762*38).</p>\n<p>Since the houses all get the same weight the contribution of each H1 line to the score should be 1/(4*1762*38) = 0.000003733</p>\n<p>I was therefore rather surprised to find out that my score for a&nbsp;<span style=\"line-height: 1.4\">submission which had only one appliance turned on for 92 minutes on H1 was 0.00068 better than the all appliances off benchmark.</span></p>\n<p><span style=\"line-height: 1.4\">I expected the maximum improvement to be 92*0.000003733 = 0.00034 and my score improved by twice that much.</span></p>\n<p>What am I missing?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28300",
      "postDate": "08/06/2013 16:29:54",
      "content": "<p>Half of the leaderboard is a public fold and half is a private fold. You are seeing your Hamming Loss only within the public fold. &nbsp;Since the split is approximately 50%, the contribution is about&nbsp;2/(4*1762*38)*92 = 0.000687, right in line with what you are seeing.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28304",
      "postDate": "08/06/2013 17:50:32",
      "content": "<p>I assume that by &quot;private fold&quot; you are referring to the testing periods provided in the data that are not included in the sample submission. &nbsp;Is that right?</p>\n<p><span style=\"line-height: 1.4\">If I understand the error messages I got correctly, it is not even possible to submit&nbsp;</span>labels for.those time periods so it is not clear to me how or why they would affect the score. &nbsp;If you are applying a factor of&nbsp;approximately 2 to the score formula, can you tell us what that factor is exactly or how to calculate it?</p>\n<p>Is the factor affected by the number of lines for each house in the private fold and if so how is it affected?</p>\n<p>If you prefer not to divulge the exact factor, can you at least confirm that the same multiplicative factor is applied to all the scores?</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28306",
      "postDate": "08/06/2013 17:58:37",
      "content": "<p>The private fold is included in the sample submission; you just don't get to see your error on that part until the end. If you saw your error on 100% of the samples, you would be able to overfit the model and/or interrogate the leaderboard for answers.</p>\n<p>We are not applying a factor of two. I just wrote it that way because 1/(1762/2) = 2/1762. In other words, the score you are getting back is only the score on 1/2 of the samples. &nbsp;The sample submission covers all the times, appliances, and houses you should predict. No more, no fewer, no sparse predictions, no multiplicative factors.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28313",
      "postDate": "08/06/2013 19:11:00",
      "content": "<p>Would it be correct to assume that the division between the private and public fold will be fixed throughout the competition?</p>\n<p>Can you tell us how many samples are being scored for each house?</p>\n<p>Are houses still weighted equally in the public score (as suggested by the evaluation function) or are the weights only approximately 1/4 because the ratio of public to private samples numbers is not the same across houses?</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28314",
      "postDate": "08/06/2013 19:44:59",
      "content": "<p>Yes, the division is fixed.</p>\n<p>It's pretty close to 50% (there is always a note above the leaderboard that tells you the split). We don't say which rows are public and which are private.</p>\n<p>Houses are weighted equally 1/4, exactly how the formula says.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28362",
      "postDate": "08/07/2013 13:40:08",
      "content": "<p>William:</p>\n<p>A. You mentioned that the private fold is included in the sample submission. &nbsp;Does that mean that test periods in the data that are not included in the sample submission are irrelevant to the competition?</p>\n<p>B. The factor of 2 that you mentioned still does not explain the differences between two of my submissions (names end in &quot;e&quot; and &quot;f&quot;). The only difference between them is one device, if the calculation you suggested is correct it would seem to indicate that in the backend solution this specific device was labeled as &quot;on&quot; for some but not all of the minutes which I marked as on. &nbsp;This seems very unlikely. &nbsp;Can you please shed some light on this issue or pass it on to someone who can explain it?</p>\n<p>C. I still believe that the device was on for the whole time and is mislabeled in the backend solution. &nbsp;I prefer to discuss the details privately with one of the admins. &nbsp;How can I contact the admins directly?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28773",
      "postDate": "08/15/2013 23:32:02",
      "content": "<p>I'm sorry but I still don't understand where the benchmark comes from. I keep getting 0.052 instead of 0.079.</p>\n<p>Here is how I calculated it, D is the sample size for each houe, I use also in the place of the XOR &nbsp;because I understand there will be 1 mistake per sample if I predict all 0.</p>\n<p>&nbsp;</p>\n<p>&gt; D1=1762 <br>&gt; D2=1831<br>&gt; D3=1238<br>&gt; D4=948<br>&gt; hamming1=(2/D1)*(D1/38)<br>&gt; hamming2=(2/D2)*(D2/38)<br>&gt; hamming3=(2/D3)*(D3/38)<br>&gt; hamming4=(2/D4)*(D4/38)<br>&gt; mean(hamming1,hamming2,hamming3,hamming4)<br>[1] 0.05263158</p>\n<p>&nbsp;</p>\n<p>Is anyone able to help me figure out what is wrong? Thank you!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28779",
      "postDate": "08/16/2013 00:46:41",
      "content": "<p>The problem likely comes from the number of apparatus, which is not a fixed 38.</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28783",
      "postDate": "08/16/2013 01:42:47",
      "content": "<p>Thanks Jay. Yeah I tough of that, still the houses have respectively 38,38,37,36 unique appliances, using them barely changes the score.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28820",
      "postDate": "08/16/2013 14:24:19",
      "content": "<p>Tiago,</p>\n<p>When I followed William's explanation of the public and private folds and used D numbers accordingly things started to make sense and the numbers work out. &nbsp;You might need to use some of your submissions to figure out the size of the public fold for each house because they are not exactly half of the total. &nbsp;I take this as part of the challenge.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 28300,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "08/06/2013 16:29:54",
      "content": "<p>Half of the leaderboard is a public fold and half is a private fold. You are seeing your Hamming Loss only within the public fold. &nbsp;Since the split is approximately 50%, the contribution is about&nbsp;2/(4*1762*38)*92 = 0.000687, right in line with what you are seeing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28304,
      "author_name": "noamtene",
      "author_url": "",
      "post_date": "08/06/2013 17:50:32",
      "content": "<p>I assume that by &quot;private fold&quot; you are referring to the testing periods provided in the data that are not included in the sample submission. &nbsp;Is that right?</p>\n<p><span style=\"line-height: 1.4\">If I understand the error messages I got correctly, it is not even possible to submit&nbsp;</span>labels for.those time periods so it is not clear to me how or why they would affect the score. &nbsp;If you are applying a factor of&nbsp;approximately 2 to the score formula, can you tell us what that factor is exactly or how to calculate it?</p>\n<p>Is the factor affected by the number of lines for each house in the private fold and if so how is it affected?</p>\n<p>If you prefer not to divulge the exact factor, can you at least confirm that the same multiplicative factor is applied to all the scores?</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28306,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "08/06/2013 17:58:37",
      "content": "<p>The private fold is included in the sample submission; you just don't get to see your error on that part until the end. If you saw your error on 100% of the samples, you would be able to overfit the model and/or interrogate the leaderboard for answers.</p>\n<p>We are not applying a factor of two. I just wrote it that way because 1/(1762/2) = 2/1762. In other words, the score you are getting back is only the score on 1/2 of the samples. &nbsp;The sample submission covers all the times, appliances, and houses you should predict. No more, no fewer, no sparse predictions, no multiplicative factors.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28313,
      "author_name": "noamtene",
      "author_url": "",
      "post_date": "08/06/2013 19:11:00",
      "content": "<p>Would it be correct to assume that the division between the private and public fold will be fixed throughout the competition?</p>\n<p>Can you tell us how many samples are being scored for each house?</p>\n<p>Are houses still weighted equally in the public score (as suggested by the evaluation function) or are the weights only approximately 1/4 because the ratio of public to private samples numbers is not the same across houses?</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28314,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "08/06/2013 19:44:59",
      "content": "<p>Yes, the division is fixed.</p>\n<p>It's pretty close to 50% (there is always a note above the leaderboard that tells you the split). We don't say which rows are public and which are private.</p>\n<p>Houses are weighted equally 1/4, exactly how the formula says.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28362,
      "author_name": "noamtene",
      "author_url": "",
      "post_date": "08/07/2013 13:40:08",
      "content": "<p>William:</p>\n<p>A. You mentioned that the private fold is included in the sample submission. &nbsp;Does that mean that test periods in the data that are not included in the sample submission are irrelevant to the competition?</p>\n<p>B. The factor of 2 that you mentioned still does not explain the differences between two of my submissions (names end in &quot;e&quot; and &quot;f&quot;). The only difference between them is one device, if the calculation you suggested is correct it would seem to indicate that in the backend solution this specific device was labeled as &quot;on&quot; for some but not all of the minutes which I marked as on. &nbsp;This seems very unlikely. &nbsp;Can you please shed some light on this issue or pass it on to someone who can explain it?</p>\n<p>C. I still believe that the device was on for the whole time and is mislabeled in the backend solution. &nbsp;I prefer to discuss the details privately with one of the admins. &nbsp;How can I contact the admins directly?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28773,
      "author_name": "tiagozortea",
      "author_url": "",
      "post_date": "08/15/2013 23:32:02",
      "content": "<p>I'm sorry but I still don't understand where the benchmark comes from. I keep getting 0.052 instead of 0.079.</p>\n<p>Here is how I calculated it, D is the sample size for each houe, I use also in the place of the XOR &nbsp;because I understand there will be 1 mistake per sample if I predict all 0.</p>\n<p>&nbsp;</p>\n<p>&gt; D1=1762 <br>&gt; D2=1831<br>&gt; D3=1238<br>&gt; D4=948<br>&gt; hamming1=(2/D1)*(D1/38)<br>&gt; hamming2=(2/D2)*(D2/38)<br>&gt; hamming3=(2/D3)*(D3/38)<br>&gt; hamming4=(2/D4)*(D4/38)<br>&gt; mean(hamming1,hamming2,hamming3,hamming4)<br>[1] 0.05263158</p>\n<p>&nbsp;</p>\n<p>Is anyone able to help me figure out what is wrong? Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28779,
      "author_name": "jay670790",
      "author_url": "",
      "post_date": "08/16/2013 00:46:41",
      "content": "<p>The problem likely comes from the number of apparatus, which is not a fixed 38.</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28783,
      "author_name": "tiagozortea",
      "author_url": "",
      "post_date": "08/16/2013 01:42:47",
      "content": "<p>Thanks Jay. Yeah I tough of that, still the houses have respectively 38,38,37,36 unique appliances, using them barely changes the score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28820,
      "author_name": "noamtene",
      "author_url": "",
      "post_date": "08/16/2013 14:24:19",
      "content": "<p>Tiago,</p>\n<p>When I followed William's explanation of the public and private folds and used D numbers accordingly things started to make sense and the numbers work out. &nbsp;You might need to use some of your submissions to figure out the size of the public fold for each house because they are not exactly half of the total. &nbsp;I take this as part of the challenge.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "28297": "",
    "28300": "",
    "28304": "",
    "28306": "",
    "28313": "",
    "28314": "",
    "28362": "",
    "28773": "",
    "28779": "",
    "28783": "",
    "28820": ""
  },
  "source": "meta"
}