{
  "id": 5111,
  "title": "Submission Format",
  "url": "/competitions/belkin-energy-disaggregation-competition/discussion/5111",
  "author_name": "",
  "post_date": "2013-07-16T06:29:13.523Z",
  "votes": null,
  "comment_count": 10,
  "views": 3189,
  "content": "<p>Can you clarify if my understanding is correct:</p>\r\n<p>1. For each house, list each of the timestamps when any of the appliances is turned on as well as when it is turned off. Is this just another way to represent the information present in AllTaggingInfo in the training sets, for the testing sets?</p>\r\nId,House,Appliance,TimeStamp,Predicted 1,H1,30,1334300400,0 2,H1,29,1334300400,0 3,H1,15,1334300400,0\r\n<p>2. The list of appliances that we are required to detect is selected from the training data for that house (AllTaggingInfo)&nbsp;</p>\r\n<p>Do we also need to list the appliances if it is never turned on (as shown in the sample submission)? What timestamp should</p>\r\n<p>we use for that?</p>\r\n<p>In addition, can you please provide some pointers on how to do event detection? I am a newbie and only have experience with classification.</p>\r\n<p></p>",
  "messages": [
    {
      "id": "27263",
      "postDate": "07/16/2013 06:29:13",
      "content": "<p>Can you clarify if my understanding is correct:</p>\r\n<p>1. For each house, list each of the timestamps when any of the appliances is turned on as well as when it is turned off. Is this just another way to represent the information present in AllTaggingInfo in the training sets, for the testing sets?</p>\r\nId,House,Appliance,TimeStamp,Predicted 1,H1,30,1334300400,0 2,H1,29,1334300400,0 3,H1,15,1334300400,0\r\n<p>2. The list of appliances that we are required to detect is selected from the training data for that house (AllTaggingInfo)&nbsp;</p>\r\n<p>Do we also need to list the appliances if it is never turned on (as shown in the sample submission)? What timestamp should</p>\r\n<p>we use for that?</p>\r\n<p>In addition, can you please provide some pointers on how to do event detection? I am a newbie and only have experience with classification.</p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27273",
      "postDate": "07/16/2013 13:26:54",
      "content": "<p>As I undertand, you're right with 1. Then, the list of appliances and timestamps required to predict are present in the Sample Submission file, and we need to predict for all of them even if one appliance is never turned on in the test data. However, all\r\n appliances present in the Sample Submission file were turned on at least once in the training data.</p>\r\n<p>The event detection is part of the competition. One approach could be to build a classifier to predict for each timestamp 1 if there is an event, and 0 if there is no event.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27596",
      "postDate": "07/24/2013 14:48:54",
      "content": "<p>I'd like some further clarification on the submission format:</p>\n<p>Column 1 - ID - is this just a running number for each row?</p>\n<p>Column 2 - House - This should be [H1,H2,H3,H4] in that order?</p>\n<p>Column 3- &nbsp;Appliance - This should be [1-38] for each house in that order?</p>\n<p>Column 4 - TimeStamp - The unix timestamp matching the timestamp in the&nbsp;Testing_XX_XX_XXXX.mat file?</p>\n<p>Column 5 - Predicted - This is [0,1] corresponding to whether that appliance is on at that timestamp?</p>\n<p>&nbsp;</p>\n<p>Should the sample submission have a row *for each* timestamp in *all* of the 'Testing_XX_XX_XXXX.mat' files, or should the sample submission have a row for each start/stop time for each appliance i.e. 2*the number of events in all of the testing data?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27708",
      "postDate": "07/26/2013 23:56:06",
      "content": "<p>I too would like an answer to gallamine's comment above. Can anyone please clarify?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27838",
      "postDate": "07/30/2013 15:07:35",
      "content": "<p>Bump.</p>\n<p>Another question would be, should there be an event number of elements in the submission file? i.e. should it only be start/stop times?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27839",
      "postDate": "07/30/2013 15:26:18",
      "content": "<p>SampleSubmission.csv is your template for what you need to provide. &nbsp;For each house, appliance, timestamp triplet listed in that file, you should fill in the Predicted column.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27840",
      "postDate": "07/30/2013 15:37:56",
      "content": "<p>The sample submission file only has 219,580 timestamps listed, whereas the Testing_XX_XX....mat files have ~8 million unique timestamps. My confusion was coming from this discrepancy.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28308",
      "postDate": "08/06/2013 18:17:22",
      "content": "<p>So, there's no need to provide a solution for the whole&nbsp;Testing_XX_XX....mat files? Just the timestamps at the SampleSubmission?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28312",
      "postDate": "08/06/2013 18:41:53",
      "content": "<p>[quote=promeu;28308]</p>\n<p>So, there's no need to provide a solution for the whole&nbsp;Testing_XX_XX....mat files? Just the timestamps at the SampleSubmission?</p>\n<p>[/quote]</p>\n<p>Correct</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "29945",
      "postDate": "09/01/2013 17:38:12",
      "content": "<p>In the training files, the tagging info have start and end timestamps quantized to the nearest minute. From forum discussions (like the &quot;Zero-length events&quot; thread) the start timestamps are taken to mean the beginning of that minute and the end timestamps are taken to mean the end of that minute.</p>\n<p>In the submission file, there is only one timestamp. Is it intended to represent the beginning of a minute or the end of a minute? This could be pretty relevant to appliances that are turned on/off quickly (e.g. garbage disposal).</p>\n<p>Thanks in advance for any clarification!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "29948",
      "postDate": "09/01/2013 18:16:30",
      "content": "<p>To answer my own question: this distinction doesn't matter, as long as the pattern of the training data start and end timestamps is preserved. A submission &quot;timestamp&quot;&nbsp;of &quot;12:34:00 pm&quot; likely means 12:34:00.000 through 12:34:59.999. In submissions, &quot;timestamps&quot; refer to an <em>interval</em>&nbsp;not a literal instant in time (which I am used to thinking of as the meaning of &quot;timestamp&quot;).</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 27273,
      "author_name": "luistp001",
      "author_url": "",
      "post_date": "07/16/2013 13:26:54",
      "content": "<p>As I undertand, you're right with 1. Then, the list of appliances and timestamps required to predict are present in the Sample Submission file, and we need to predict for all of them even if one appliance is never turned on in the test data. However, all\r\n appliances present in the Sample Submission file were turned on at least once in the training data.</p>\r\n<p>The event detection is part of the competition. One approach could be to build a classifier to predict for each timestamp 1 if there is an event, and 0 if there is no event.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27596,
      "author_name": "gallamine",
      "author_url": "",
      "post_date": "07/24/2013 14:48:54",
      "content": "<p>I'd like some further clarification on the submission format:</p>\n<p>Column 1 - ID - is this just a running number for each row?</p>\n<p>Column 2 - House - This should be [H1,H2,H3,H4] in that order?</p>\n<p>Column 3- &nbsp;Appliance - This should be [1-38] for each house in that order?</p>\n<p>Column 4 - TimeStamp - The unix timestamp matching the timestamp in the&nbsp;Testing_XX_XX_XXXX.mat file?</p>\n<p>Column 5 - Predicted - This is [0,1] corresponding to whether that appliance is on at that timestamp?</p>\n<p>&nbsp;</p>\n<p>Should the sample submission have a row *for each* timestamp in *all* of the 'Testing_XX_XX_XXXX.mat' files, or should the sample submission have a row for each start/stop time for each appliance i.e. 2*the number of events in all of the testing data?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27708,
      "author_name": "",
      "author_url": "",
      "post_date": "07/26/2013 23:56:06",
      "content": "<p>I too would like an answer to gallamine's comment above. Can anyone please clarify?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27838,
      "author_name": "gallamine",
      "author_url": "",
      "post_date": "07/30/2013 15:07:35",
      "content": "<p>Bump.</p>\n<p>Another question would be, should there be an event number of elements in the submission file? i.e. should it only be start/stop times?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27839,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "07/30/2013 15:26:18",
      "content": "<p>SampleSubmission.csv is your template for what you need to provide. &nbsp;For each house, appliance, timestamp triplet listed in that file, you should fill in the Predicted column.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27840,
      "author_name": "gallamine",
      "author_url": "",
      "post_date": "07/30/2013 15:37:56",
      "content": "<p>The sample submission file only has 219,580 timestamps listed, whereas the Testing_XX_XX....mat files have ~8 million unique timestamps. My confusion was coming from this discrepancy.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28308,
      "author_name": "pabloromeu",
      "author_url": "",
      "post_date": "08/06/2013 18:17:22",
      "content": "<p>So, there's no need to provide a solution for the whole&nbsp;Testing_XX_XX....mat files? Just the timestamps at the SampleSubmission?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28312,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "08/06/2013 18:41:53",
      "content": "<p>[quote=promeu;28308]</p>\n<p>So, there's no need to provide a solution for the whole&nbsp;Testing_XX_XX....mat files? Just the timestamps at the SampleSubmission?</p>\n<p>[/quote]</p>\n<p>Correct</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 29945,
      "author_name": "rosnfeld",
      "author_url": "",
      "post_date": "09/01/2013 17:38:12",
      "content": "<p>In the training files, the tagging info have start and end timestamps quantized to the nearest minute. From forum discussions (like the &quot;Zero-length events&quot; thread) the start timestamps are taken to mean the beginning of that minute and the end timestamps are taken to mean the end of that minute.</p>\n<p>In the submission file, there is only one timestamp. Is it intended to represent the beginning of a minute or the end of a minute? This could be pretty relevant to appliances that are turned on/off quickly (e.g. garbage disposal).</p>\n<p>Thanks in advance for any clarification!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 29948,
      "author_name": "rosnfeld",
      "author_url": "",
      "post_date": "09/01/2013 18:16:30",
      "content": "<p>To answer my own question: this distinction doesn't matter, as long as the pattern of the training data start and end timestamps is preserved. A submission &quot;timestamp&quot;&nbsp;of &quot;12:34:00 pm&quot; likely means 12:34:00.000 through 12:34:59.999. In submissions, &quot;timestamps&quot; refer to an <em>interval</em>&nbsp;not a literal instant in time (which I am used to thinking of as the meaning of &quot;timestamp&quot;).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "27263": "",
    "27273": "",
    "27596": "",
    "27708": "",
    "27838": "",
    "27839": "",
    "27840": "",
    "28308": "",
    "28312": "",
    "29945": "",
    "29948": ""
  },
  "source": "meta"
}