{
  "id": 5169,
  "title": "one or two stage competition?",
  "url": "/competitions/belkin-energy-disaggregation-competition/discussion/5169",
  "author_name": "",
  "post_date": "2013-07-22T08:29:25.347Z",
  "votes": null,
  "comment_count": 3,
  "views": 1489,
  "content": "<p>Is this a one stage competition? By this I mean that we only get one set of data to predict and whoever does that best wins the prize.</p>\n<p>&nbsp;&nbsp;&nbsp;&nbsp; (If so why does it say above the leaderboard; &quot;This leaderboard is&nbsp;&nbsp;&nbsp; calculated on approximately 50% of the test data.<br> The final results will be based on the other 50%, so the final standings may be different&quot;?)</p>\n<p>Or is this a two stage competition whereby we have to submit our final model (or hash of it) and than get to predict the final evaluation set?</p>\n<p>&nbsp;&nbsp;&nbsp;&nbsp; (If so when do we have to submit our final models and when is the final evaluation set released?)</p>\n<p>Or is it none of the above?</p>\n<p>&nbsp;</p>\n<p>Thanks in advance!</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
  "messages": [
    {
      "id": "27502",
      "postDate": "07/22/2013 08:29:25",
      "content": "<p>Is this a one stage competition? By this I mean that we only get one set of data to predict and whoever does that best wins the prize.</p>\n<p>&nbsp;&nbsp;&nbsp;&nbsp; (If so why does it say above the leaderboard; &quot;This leaderboard is&nbsp;&nbsp;&nbsp; calculated on approximately 50% of the test data.<br> The final results will be based on the other 50%, so the final standings may be different&quot;?)</p>\n<p>Or is this a two stage competition whereby we have to submit our final model (or hash of it) and than get to predict the final evaluation set?</p>\n<p>&nbsp;&nbsp;&nbsp;&nbsp; (If so when do we have to submit our final models and when is the final evaluation set released?)</p>\n<p>Or is it none of the above?</p>\n<p>&nbsp;</p>\n<p>Thanks in advance!</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27508",
      "postDate": "07/22/2013 13:50:18",
      "content": "<p>This is one stage. &nbsp;You see your error on a random and fixed 50% of the test set. The final ranking is based on the other 50%.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27537",
      "postDate": "07/23/2013 07:18:45",
      "content": "<p>Ah, thanks William. I was interpreting &quot;is calculated on 50% of the test data&quot; as if we <strong>didn't</strong> receive<br><strong>all</strong> the test data yet.</p>\n<p>That brings me to another question:</p>\n<p>How is the distribution of the test data randomized over the two parts (the part used for calculating the score on leaderboard and the final evaluation set)? Is it on row level? (i.e. each row has a 50% chance to be in the final evaluation part) . Or is it done at another level e.g. on the level of houses, appliances, time periods or combinations of those? (e.g. the history of appliance A in house H1 from time stamp x till time stamp y is either entirely in the final evaluation set or it is not)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27551",
      "postDate": "07/23/2013 15:27:56",
      "content": "<p>It is randomized at the row level.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 27508,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "07/22/2013 13:50:18",
      "content": "<p>This is one stage. &nbsp;You see your error on a random and fixed 50% of the test set. The final ranking is based on the other 50%.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27537,
      "author_name": "julesvanligtenberg",
      "author_url": "",
      "post_date": "07/23/2013 07:18:45",
      "content": "<p>Ah, thanks William. I was interpreting &quot;is calculated on 50% of the test data&quot; as if we <strong>didn't</strong> receive<br><strong>all</strong> the test data yet.</p>\n<p>That brings me to another question:</p>\n<p>How is the distribution of the test data randomized over the two parts (the part used for calculating the score on leaderboard and the final evaluation set)? Is it on row level? (i.e. each row has a 50% chance to be in the final evaluation part) . Or is it done at another level e.g. on the level of houses, appliances, time periods or combinations of those? (e.g. the history of appliance A in house H1 from time stamp x till time stamp y is either entirely in the final evaluation set or it is not)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27551,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "07/23/2013 15:27:56",
      "content": "<p>It is randomized at the row level.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "27502": "",
    "27508": "",
    "27537": "",
    "27551": ""
  },
  "source": "meta"
}