{
  "id": 3825,
  "title": "Model Submission Instructions",
  "url": "/competitions/flight/discussion/3825",
  "author_name": "",
  "post_date": "2013-02-13T06:33:42.873Z",
  "votes": null,
  "comment_count": 7,
  "views": 6143,
  "content": "<p><span style=\"font-size:14px; line-height:1.4em\">As a reminder, you have until 6:00pm UTC on February 15 to upload a hash or archive of your final model to the site. Basically, the final model should consist of a single archive (or SHA-256 hash thereof) that\r\n contains all the data and code necessary to make predictions on the final evaluation set (which will be released on March 4). At that point, you will then have a week to run your model on the final evaluation set and submit these predictions.</span></p>\r\n<p>To submit your model to the site, go to the &quot;My Submissions&quot; section of this competition and then click &quot;add model hash&quot; or &quot;add attachment&quot; as appropriate on your best-scoring public submission.</p>\r\n<p>Also, we've started drafting a <a href=\"https://www.kaggle.com/wiki/ModelSubmissionBestPractices\">\r\nbest practices guide for model submissions</a>. We don't expect models submitted for this competition to conform precisely to these best practices, but wanted to go ahead and release them to get early feedback and as rough guidelines.</p>\r\n<p>Please let us know in this forum if you have any questions/comments on the model submission process or the &quot;best practices&quot; guide.</p>\r\n<p>Thanks for your participation in this contest so far, and good luck as it enters its final stages!</p>",
  "messages": [
    {
      "id": "20402",
      "postDate": "02/13/2013 06:33:42",
      "content": "<p><span style=\"font-size:14px; line-height:1.4em\">As a reminder, you have until 6:00pm UTC on February 15 to upload a hash or archive of your final model to the site. Basically, the final model should consist of a single archive (or SHA-256 hash thereof) that\r\n contains all the data and code necessary to make predictions on the final evaluation set (which will be released on March 4). At that point, you will then have a week to run your model on the final evaluation set and submit these predictions.</span></p>\r\n<p>To submit your model to the site, go to the &quot;My Submissions&quot; section of this competition and then click &quot;add model hash&quot; or &quot;add attachment&quot; as appropriate on your best-scoring public submission.</p>\r\n<p>Also, we've started drafting a <a href=\"https://www.kaggle.com/wiki/ModelSubmissionBestPractices\">\r\nbest practices guide for model submissions</a>. We don't expect models submitted for this competition to conform precisely to these best practices, but wanted to go ahead and release them to get early feedback and as rough guidelines.</p>\r\n<p>Please let us know in this forum if you have any questions/comments on the model submission process or the &quot;best practices&quot; guide.</p>\r\n<p>Thanks for your participation in this contest so far, and good luck as it enters its final stages!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "20410",
      "postDate": "02/13/2013 09:07:01",
      "content": "<p><span style=\"line-height:1.4em\">Will the general structure of the test days be the same as the PublicLeaderboardSet archive? Can You tell me a bit about the format of the file that contains the cutoffs for days to be predicted? Will it be the same as the\r\n days.csv as in PublicLeaderboardSet?&nbsp;</span><span style=\"line-height:1.4em\">Would there be a&nbsp;test_flights.csv in the folder of the test days just like there were in the PublicLeaderboardSet?</span></p>\r\n<p><span style=\"line-height:1.4em\">Are we supposed to provide preprocessing code for the training set or just include the data generated that we train our model on?\r\n</span></p>\r\n<p><span style=\"line-height:1.4em\">For the yet unknown test set what kind of assumptions could be made? For example, my code would fail - and many other's as well I presume - if there are no flighthistoryevents.csv table (or it is called flighthistory_events.csv\r\n or whatever)? How should one overcome this problem or uncertainty?&nbsp;</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "20432",
      "postDate": "02/13/2013 17:23:46",
      "content": "<p>Ben,</p>\r\n<p>According to the countdown timer, the deadline would be 0AM<span>&nbsp;UTC on February 15 rather than 6PM, right?</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "20433",
      "postDate": "02/13/2013 18:30:07",
      "content": "<p>The rules as posted at&nbsp;https://www.gequest.com/c/flight/details/submission-instructions state:</p>\r\n<p><span>Your final model must contain all code and parameter settings necessary to evaluate your models on new data, and include a README file with instructions on how to do so.</span></p>\r\n<p>They do not say that the software used to generate the parameters (i.e., for training the model) need to be included.</p>\r\n<p>The top post, however, says the we must upload an archive or hash, &quot;...<span>&nbsp;that contains all the code necessary to train your model and make predictions on the final evaluation set.&quot;</span></p>\r\n<p><span>This is ambiguous. Does it mean</span></p>\r\n<p><span>1) Upload the code to train your model on the training data. And upload the code to make predictions on the final evaluation set.</span></p>\r\n<p><span>or does it mean</span></p>\r\n<p><span>2)&nbsp;Upload the code to train your model on the final evaluation set and make predictions on the final evaluation set.</span></p>\r\n<p><span>In other words, is this a NEW requirement that we must also upload the software to generate our parameters from the training data? In my case, that's a much larger collection of code. Or do we just upload the code and (pre-defined) tables of parameters\r\n that can apply to the new test set?</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "20440",
      "postDate": "02/13/2013 20:59:39",
      "content": "<p>[quote=Guocong Song;20432]</p>\r\n<p>Ben,</p>\r\n<p>According to the countdown timer, the deadline would be 0AM<span>&nbsp;UTC on February 15 rather than 6PM, right?</span></p>\r\n<p>[/quote]Thanks for pointing this out, the deadline has been updated to correspond to the model submission deadline (6pm UTC on February 15).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "20447",
      "postDate": "02/13/2013 23:04:31",
      "content": "<p>[quote=ivo;20410]</p>\r\n<p><span style=\"line-height:1.4em\">Will the general structure of the test days be the same as the PublicLeaderboardSet archive? Can You tell me a bit about the format of the file that contains the cutoffs for days to be predicted? Will it be the same as the\r\n days.csv as in PublicLeaderboardSet?&nbsp;</span><span style=\"line-height:1.4em\">Would there be a&nbsp;test_flights.csv in the folder of the test days just like there were in the PublicLeaderboardSet?</span></p>\r\n<p>[/quote]The structure of the test days will be the same as that of the PublicLeaderboardSet.</p>\r\n<p><span>[quote=ivo;20410]Are we supposed to provide preprocessing code for the training set or just include the data generated that we train our model on?[/quote]If it is straightforward for you to include this, you should go ahead and include it. However,\r\n this is not mandatory.</span></p>\r\n<p><span>[quote=ivo;20410]For the yet unknown test set what kind of assumptions could be made? For example, my code would fail - and many other's as well I presume - if there are no flighthistoryevents.csv table (or it is called flighthistory_events.csv or\r\n whatever)? How should one overcome this problem or uncertainty?[/quote]Bug fixes to correct for data formatting / IO issues that arise will be permitted (you will need to make the case that this isn't a meangingful change to the model if this arises, simply\r\n fixing a data formatting bug).</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "20448",
      "postDate": "02/13/2013 23:12:56",
      "content": "<p>[quote=Michael 7402;20433]</p>\r\n<p>The rules as posted at&nbsp;https://www.gequest.com/c/flight/details/submission-instructions state:</p>\r\n<p><span>Your final model must contain all code and parameter settings necessary to evaluate your models on new data, and include a README file with instructions on how to do so.</span></p>\r\n<p>They do not say that the software used to generate the parameters (i.e., for training the model) need to be included.</p>\r\n<p>The top post, however, says the we must upload an archive or hash, &quot;...<span>&nbsp;that contains all the code necessary to train your model and make predictions on the final evaluation set.&quot;</span></p>\r\n<p><span>This is ambiguous. Does it mean</span></p>\r\n<p><span>1) Upload the code to train your model on the training data. And upload the code to make predictions on the final evaluation set.</span></p>\r\n<p><span>or does it mean</span></p>\r\n<p><span>2)&nbsp;Upload the code to train your model on the final evaluation set and make predictions on the final evaluation set.</span></p>\r\n<p><span>In other words, is this a NEW requirement that we must also upload the software to generate our parameters from the training data? In my case, that's a much larger collection of code. Or do we just upload the code and (pre-defined) tables of parameters\r\n that can apply to the new test set?</span></p>\r\n<p>[/quote]Sorry for the confusion - there are no new requirements. The submission instructions you linked have it correct, and I've edited my original post accordingly. It is not required to include the code to train your model on the training data in the\r\n model submission, only the model itself and the code to make predictions on a new day. However, if it straightforward to including the training code as well, I recommend doing so (this could make things easier to debug in the event that something doesn't work).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "20506",
      "postDate": "02/15/2013 18:07:59",
      "content": "<p>sigh... According to the uber time clock in the house the time was 5:58pm UTC when I went to find my submissions to attach a last code archive containing a little polish and a few tweaks. &nbsp;I was unable to access 'My Submissions' to attach the final code\r\n archive. &nbsp;Did the model submission deadline end early? &nbsp;Did anyone else run into this? &nbsp;Luckily I had attached a rougher version earlier...&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>ps. &nbsp;The uber time clock is made by EndRun Technologies and is accurate to 10 micro seconds.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 20410,
      "author_name": "iamivo",
      "author_url": "",
      "post_date": "02/13/2013 09:07:01",
      "content": "<p><span style=\"line-height:1.4em\">Will the general structure of the test days be the same as the PublicLeaderboardSet archive? Can You tell me a bit about the format of the file that contains the cutoffs for days to be predicted? Will it be the same as the\r\n days.csv as in PublicLeaderboardSet?&nbsp;</span><span style=\"line-height:1.4em\">Would there be a&nbsp;test_flights.csv in the folder of the test days just like there were in the PublicLeaderboardSet?</span></p>\r\n<p><span style=\"line-height:1.4em\">Are we supposed to provide preprocessing code for the training set or just include the data generated that we train our model on?\r\n</span></p>\r\n<p><span style=\"line-height:1.4em\">For the yet unknown test set what kind of assumptions could be made? For example, my code would fail - and many other's as well I presume - if there are no flighthistoryevents.csv table (or it is called flighthistory_events.csv\r\n or whatever)? How should one overcome this problem or uncertainty?&nbsp;</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 20432,
      "author_name": "songgc",
      "author_url": "",
      "post_date": "02/13/2013 17:23:46",
      "content": "<p>Ben,</p>\r\n<p>According to the countdown timer, the deadline would be 0AM<span>&nbsp;UTC on February 15 rather than 6PM, right?</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 20433,
      "author_name": "budiansky",
      "author_url": "",
      "post_date": "02/13/2013 18:30:07",
      "content": "<p>The rules as posted at&nbsp;https://www.gequest.com/c/flight/details/submission-instructions state:</p>\r\n<p><span>Your final model must contain all code and parameter settings necessary to evaluate your models on new data, and include a README file with instructions on how to do so.</span></p>\r\n<p>They do not say that the software used to generate the parameters (i.e., for training the model) need to be included.</p>\r\n<p>The top post, however, says the we must upload an archive or hash, &quot;...<span>&nbsp;that contains all the code necessary to train your model and make predictions on the final evaluation set.&quot;</span></p>\r\n<p><span>This is ambiguous. Does it mean</span></p>\r\n<p><span>1) Upload the code to train your model on the training data. And upload the code to make predictions on the final evaluation set.</span></p>\r\n<p><span>or does it mean</span></p>\r\n<p><span>2)&nbsp;Upload the code to train your model on the final evaluation set and make predictions on the final evaluation set.</span></p>\r\n<p><span>In other words, is this a NEW requirement that we must also upload the software to generate our parameters from the training data? In my case, that's a much larger collection of code. Or do we just upload the code and (pre-defined) tables of parameters\r\n that can apply to the new test set?</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 20440,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "02/13/2013 20:59:39",
      "content": "<p>[quote=Guocong Song;20432]</p>\r\n<p>Ben,</p>\r\n<p>According to the countdown timer, the deadline would be 0AM<span>&nbsp;UTC on February 15 rather than 6PM, right?</span></p>\r\n<p>[/quote]Thanks for pointing this out, the deadline has been updated to correspond to the model submission deadline (6pm UTC on February 15).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 20447,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "02/13/2013 23:04:31",
      "content": "<p>[quote=ivo;20410]</p>\r\n<p><span style=\"line-height:1.4em\">Will the general structure of the test days be the same as the PublicLeaderboardSet archive? Can You tell me a bit about the format of the file that contains the cutoffs for days to be predicted? Will it be the same as the\r\n days.csv as in PublicLeaderboardSet?&nbsp;</span><span style=\"line-height:1.4em\">Would there be a&nbsp;test_flights.csv in the folder of the test days just like there were in the PublicLeaderboardSet?</span></p>\r\n<p>[/quote]The structure of the test days will be the same as that of the PublicLeaderboardSet.</p>\r\n<p><span>[quote=ivo;20410]Are we supposed to provide preprocessing code for the training set or just include the data generated that we train our model on?[/quote]If it is straightforward for you to include this, you should go ahead and include it. However,\r\n this is not mandatory.</span></p>\r\n<p><span>[quote=ivo;20410]For the yet unknown test set what kind of assumptions could be made? For example, my code would fail - and many other's as well I presume - if there are no flighthistoryevents.csv table (or it is called flighthistory_events.csv or\r\n whatever)? How should one overcome this problem or uncertainty?[/quote]Bug fixes to correct for data formatting / IO issues that arise will be permitted (you will need to make the case that this isn't a meangingful change to the model if this arises, simply\r\n fixing a data formatting bug).</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 20448,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "02/13/2013 23:12:56",
      "content": "<p>[quote=Michael 7402;20433]</p>\r\n<p>The rules as posted at&nbsp;https://www.gequest.com/c/flight/details/submission-instructions state:</p>\r\n<p><span>Your final model must contain all code and parameter settings necessary to evaluate your models on new data, and include a README file with instructions on how to do so.</span></p>\r\n<p>They do not say that the software used to generate the parameters (i.e., for training the model) need to be included.</p>\r\n<p>The top post, however, says the we must upload an archive or hash, &quot;...<span>&nbsp;that contains all the code necessary to train your model and make predictions on the final evaluation set.&quot;</span></p>\r\n<p><span>This is ambiguous. Does it mean</span></p>\r\n<p><span>1) Upload the code to train your model on the training data. And upload the code to make predictions on the final evaluation set.</span></p>\r\n<p><span>or does it mean</span></p>\r\n<p><span>2)&nbsp;Upload the code to train your model on the final evaluation set and make predictions on the final evaluation set.</span></p>\r\n<p><span>In other words, is this a NEW requirement that we must also upload the software to generate our parameters from the training data? In my case, that's a much larger collection of code. Or do we just upload the code and (pre-defined) tables of parameters\r\n that can apply to the new test set?</span></p>\r\n<p>[/quote]Sorry for the confusion - there are no new requirements. The submission instructions you linked have it correct, and I've edited my original post accordingly. It is not required to include the code to train your model on the training data in the\r\n model submission, only the model itself and the code to make predictions on a new day. However, if it straightforward to including the training code as well, I recommend doing so (this could make things easier to debug in the event that something doesn't work).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 20506,
      "author_name": "hurlburt",
      "author_url": "",
      "post_date": "02/15/2013 18:07:59",
      "content": "<p>sigh... According to the uber time clock in the house the time was 5:58pm UTC when I went to find my submissions to attach a last code archive containing a little polish and a few tweaks. &nbsp;I was unable to access 'My Submissions' to attach the final code\r\n archive. &nbsp;Did the model submission deadline end early? &nbsp;Did anyone else run into this? &nbsp;Luckily I had attached a rougher version earlier...&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>ps. &nbsp;The uber time clock is made by EndRun Technologies and is accurate to 10 micro seconds.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "20402": "",
    "20410": "",
    "20432": "",
    "20433": "",
    "20440": "",
    "20447": "",
    "20448": "",
    "20506": ""
  },
  "source": "meta"
}