{
  "id": 3220,
  "title": "Contradiction in rules (for final model)",
  "url": "/competitions/flight/discussion/3220",
  "author_name": "",
  "post_date": "2012-11-30T19:26:28.253Z",
  "votes": 2,
  "comment_count": 29,
  "views": 24777,
  "content": "<p>Hello</p>\r\n<p>I am seeing a contradiction in the rules for what is allowable for final model.</p>\r\n<p>At this link: http://www.gequest.com/c/flight/details/submission-instructions it states:</p>\r\n<p>[quote]Your model must be structured so that it makes each test day's predictions\r\n<strong>based on no information in the final evaluation test data other than the information from that day</strong>, which will be in an appropriately named folder.[/quote]</p>\r\n<p>So for prediction of Feb 20, it <span style=\"text-decoration:underline\">is not allowed</span> to use data from Feb 15 thru Feb 19.</p>\r\n<p>However in http://www.gequest.com/c/flight/data and also in https://www.gequest.com/wiki/FlightQuest.Data it states:</p>\r\n<p>[quote]For each day in the test period (first for the public leaderboard, and later for the final evaluation), we will select a random time (uniformly chosen between 9am EST and 9pm EST) and select all of the flights in the air at that cutoff time. You will\r\n be provided with relevant data for each day that would be available at the chosen cutoff time.\r\n<strong>Predictions for each flight on a given day can not reference any data related to future dates in the evaluation data set.</strong>[/quote]</p>\r\n<p>So for prediction of Feb 20, it <span style=\"text-decoration:underline\">is allowed</span> to use data from Feb 15 thru 19, (but not Feb 21 and beyond.)</p>\r\n<hr>\r\n<p>So for the final evaluation set, which is correct?&nbsp;&nbsp; Is prior-dated data allowed to be considered when making a prediction or not?</p>\r\n<p>Thanks</p>\r\n<p>[edit: spelling]</p>",
  "messages": [
    {
      "id": "17281",
      "postDate": "11/30/2012 19:26:28",
      "content": "<p>Hello</p>\r\n<p>I am seeing a contradiction in the rules for what is allowable for final model.</p>\r\n<p>At this link: http://www.gequest.com/c/flight/details/submission-instructions it states:</p>\r\n<p>[quote]Your model must be structured so that it makes each test day's predictions\r\n<strong>based on no information in the final evaluation test data other than the information from that day</strong>, which will be in an appropriately named folder.[/quote]</p>\r\n<p>So for prediction of Feb 20, it <span style=\"text-decoration:underline\">is not allowed</span> to use data from Feb 15 thru Feb 19.</p>\r\n<p>However in http://www.gequest.com/c/flight/data and also in https://www.gequest.com/wiki/FlightQuest.Data it states:</p>\r\n<p>[quote]For each day in the test period (first for the public leaderboard, and later for the final evaluation), we will select a random time (uniformly chosen between 9am EST and 9pm EST) and select all of the flights in the air at that cutoff time. You will\r\n be provided with relevant data for each day that would be available at the chosen cutoff time.\r\n<strong>Predictions for each flight on a given day can not reference any data related to future dates in the evaluation data set.</strong>[/quote]</p>\r\n<p>So for prediction of Feb 20, it <span style=\"text-decoration:underline\">is allowed</span> to use data from Feb 15 thru 19, (but not Feb 21 and beyond.)</p>\r\n<hr>\r\n<p>So for the final evaluation set, which is correct?&nbsp;&nbsp; Is prior-dated data allowed to be considered when making a prediction or not?</p>\r\n<p>Thanks</p>\r\n<p>[edit: spelling]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17284",
      "postDate": "11/30/2012 19:36:42",
      "content": "<p>Additionally, it appears that this wording even implies that the model should &quot;start clean&quot; on each test. IOW, it would be an implicit violation if your model has ever been trained on data from the future from the standpoint of the randomly chosen test day.\r\n</p>\r\n<p>This is a condition that may prove nearly impossible to implement since it means you must retrain your model fresh for each test. Any learned information in the model would need to explicitly come from past days only. Since it's impossible to identify the\r\n source of learned data in most models, we must either re-train on every test, or assume that the test period will be chosen from a hidden set of dates we won't have access to in the training set.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17296",
      "postDate": "11/30/2012 21:35:04",
      "content": "<p>Our apologies -- the wiki and data page were wrong; the submission instructions page was correct. Now they are all correct.</p>\r\n<p>Your model must be structured to make final test data set predictions based on no days in the final test data set other than the appropriate one (none prior or later are allowed).\r\n</p>\r\n<p>This rule only applies to the final test data set. Note that some of the training data will be from a period later than the public leaderboard test data set. It is fine to use this data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17317",
      "postDate": "12/01/2012 11:31:33",
      "content": "<p>&quot;For each day in the test period (first for the public leaderboard, and later for the final evaluation), we will select a random time (uniformly chosen between 9am EST and 9pm EST) and select all of the flights in the air at that cutoff time. You will be\r\n provided with relevant data for each day that would be available at the chosen cutoff time.&quot;</p>\r\n<p><strong>&quot;Your model must be structured so that it makes each test day's final test data set predictions based on no information in the final evaluation test data other than the information from that day, which will be in an appropriately named folder.&nbsp;</strong>(Reworded\r\n for clarification on 11/30/2012. See&nbsp;<a href=\"https://www.gequest.com/c/flight/forums/t/3220/contradiction-in-rules-for-final-model\">forum</a>&nbsp;for explanation.&quot;</p>\r\n<p>It would be useful to people who are new to the flight data domain if the above could be explained with clarity by the organizers. Walk it through with an example or something that assists understanding. &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17334",
      "postDate": "12/02/2012 05:33:30",
      "content": "<ol>\r\n<li>I am guessing that for each test day we can use all data (including Arrival Times Scheduled and Actual) for that day up until a random cutoff time and then we will predict the arrival times for the rest of that day of only the aircraft that are in the air\r\n at the time of the cutoff. This would mean no extra points for good predictions of how long it takes the plane to board, taxi down the runway, and get into the air.\r\n</li></ol>\r\n<p>Can anyone tell me if this sounds correct?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17376",
      "postDate": "12/04/2012 05:53:19",
      "content": "<p>Just how far does this use of time constrained data go? Given that, almost by definition, any model used will have been derived from data outside this period, including constants, structure, formulae etc.&nbsp;</p>\r\n<p>Some further clarity would be helpful to keep pedantry under control. An example could well help too.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17383",
      "postDate": "12/04/2012 17:32:58",
      "content": "<p>I'm still little bit confused about the objective of the competition.</p>\r\n<p>Let's assume I am a pilot flying, right now, towards my destination airport.</p>\r\n<p>A wanted model, such as a final model (whose prediction performance should have been measured at a particular cut-off time), will give me a fairly good estimation for my actual arrival time on that airport. So, I have now better estimation for my arrival\r\n time which differs from the original flight plan.</p>\r\n<p>Then, how can I use this knowledge to improve 'efficiency' of my flight? Do I have to increase/reduce my plane's speed to stick closer to the original plan?</p>\r\n<p>Could anyone, please, clarify the meaning of the objective to predict arrival time in terms of &quot;Make flying more efficient?&quot;</p>\r\n<p>Thanks.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17418",
      "postDate": "12/05/2012 19:00:46",
      "content": "<p>Joexjmmvhm-- that is correct.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17419",
      "postDate": "12/05/2012 19:04:23",
      "content": "<p>MarkA-- your models (parameters,etc.) will be based on the training data. This model then needs to be able to look at one day of the final test data (in isolation) and make predictions. So the sequence is:</p>\r\n<p>1) Build models based on training data, submitting to public leaderboard<br>\r\n2) Finalize code based on full training data set -- submit code -- code (which will include all model parameters, etc.) must be able to look at one day of test data in isolation and make predictions for that day<br>\r\n3) Final test data is released<br>\r\n4) Submit predictions for test data, using submitted code<br>\r\n5) Winners' code will be verified for compliance with these rules</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17420",
      "postDate": "12/05/2012 19:06:30",
      "content": "<p>Farstar-- the pilot wouldn't use these predictions directly. Instead, these predictions will feed into a broader optimization of parameters like the ones you mentioned, like speed. You can't optimize over the outcomes until you know which situations will\r\n lead to which outcomes.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18332",
      "postDate": "12/21/2012 12:34:40",
      "content": "<p>David, you wrote:</p>\r\n<blockquote>\r\n<p>1) Build models based on training data, submitting to public leaderboard<br>\r\n2) Finalize code based on full training data set -- submit code -- code (which will include all model parameters, etc.) must be able to look at one day of test data in isolation and make predictions for that day</p>\r\n</blockquote>\r\n<p>You see, when you &quot;train&quot; a model, it can actually memorise all the data it has seen. So your statement that you &quot;allow to train model on all data and do not allow to use all data before cutoff date for scoring&quot; does not make much sense.</p>\r\n<p>Let me give you an example - you are building a KNN model that actually memorises all data in it and then simply finds closest neighbours. Although you might think that you are using only current record (data from the current day) for scoring, other records\r\n from previous days are used as well because you need to calculate distances to them.</p>\r\n<p>Because you allow to train on all data, my model can memorise it all as a parameters that you allow to store.</p>\r\n<p>Let me give you another example. For this challenge I want to build this type of naive model:</p>\r\n<p>1. For each airport I canculate average delay of arrival (average difference between actual and estimated arrrival times).</p>\r\n<p>2. For each flight in the air I simply subtract estimated arrival time from the cutoff time and add average delay for the destination airport.</p>\r\n<p>So here is the question. You say that this naive model is not allowed, because to calculate it I am using average delays for airports that I calculated using all the data.To my mind this rule does not make practical sence because it contradicts with your\r\n statement that you allow to save model parameters.</p>\r\n<p>Because any parameters are stored as bytes and all the data from previous days is stored as bytes you cannot allow one thing and prohibit another.</p>\r\n<p>Do you agree?</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18333",
      "postDate": "12/21/2012 12:59:17",
      "content": "<p>Hear, hear... I was kinda wondering about the same thing (i.e. ban on using data from other days vs permission to use parameters trained on those days)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18335",
      "postDate": "12/21/2012 13:25:57",
      "content": "<p>I have been interpreting it this way. &nbsp;Any data that is available from the competition before the model submission deadline (and prior to the final evaluation set is released) is fair game for the models. &nbsp;Then, when the final data set is released, the model\r\n execution should be restricted to the one day of test data. &nbsp;So, a KNN model could contain all the training data, but should not add any of the instances from the other days in the final evaluation set to the neighborhood.</p>\r\n<p>For example, a model should produce the exact same predictions for day 4 in the final evaluation set, regardless of how many other days of data are provided in that final evaluation set.</p>\r\n<p>Is this the intent?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18339",
      "postDate": "12/21/2012 13:56:12",
      "content": "<p>Fine: let's take a linear regression model as an example. If I have estimated the parameters on e.g. all days from november and then I try to use this model taking only december 4th as input - I am still - implicitly - using the other observations, because\r\n that affects what my regression parameters look like - so technically, I am in violation of the rules. If you want to avoid this, then training data should not be used at all...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18342",
      "postDate": "12/21/2012 14:18:05",
      "content": "<p>[quote=BJG_;18335]</p>\r\n<p>I have been interpreting it this way. &nbsp;Any data that is available from the competition before the model submission deadline (and prior to the final evaluation set is released) is fair game for the models. &nbsp;Then, when the final data set is released, the model\r\n execution should be restricted to the one day of test data. &nbsp;So, a KNN model could contain all the training data, but should not add any of the instances from the other days in the final evaluation set to the neighborhood.</p>\r\n<p>For example, a model should produce the exact same predictions for day 4 in the final evaluation set, regardless of how many other days of data are provided in that final evaluation set.</p>\r\n<p>Is this the intent?</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>This intent makes no sense. You can imagine a model as a black box loaded with some universal knowledge that makes predictions. You come to this black box, open it, put your new day data into it, step back and then ask &quot;Dear box, when do you think this plane\r\n that I have just told you about, will land?&quot;. The box uses ALL of its knowledge to answer this question. Both your current data and everything it has seen before.</p>\r\n<p>Here is an example - IBM Watson machine that has beaten humans in Jeopardy game. It was so clever because it had access to HUGE amounts of text data and it could search through it in real time and find answers.</p>\r\n<p>Another example - you are playing chess with AI. You are loosing because it has millions of strategies loaded in it and it is optimising its chanses to win by predicting which strategy is the best. And you say: &quot;Stop doing that! This is not fair! Please,\r\n do not use all your strategies that you have learned from previous games. You can use data only from this game!&quot; =)&nbsp; Ha-ha, that would be funny =)</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18344",
      "postDate": "12/21/2012 14:28:51",
      "content": "<p>I'm guessing the primary intent is to avoid using the future data in the prediction. &nbsp;They do not want a model to use data from Day 14 in the final evaluation set to make predictions for Day 4.&nbsp;</p>\r\n<p>There is also a chance that the customer is risk adverse in this area and does not want to have a model whose parameters change automatically after it is deployed. &nbsp;They may plan to go through some testing process prior to deploying the model and want any\r\n changes to the model to go through a similar testing process (not learned automatically). &nbsp;</p>\r\n<p>&nbsp;I am just conjecturing though.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18350",
      "postDate": "12/21/2012 15:28:17",
      "content": "<p>It's not like Jeopardy, but it may be like the chess example.<br>\r\nI think the thing to do is to look at how this will be used in the real world. <br>\r\nThey do not want the tool to go through all of history and also include today's data and come to some analysis.<br>\r\nIt will not have access to all this data, nor will it have all this time to do analysis and arrive at an answer.</p>\r\n<p>You go to school for x number of years, and you read many books and learn from them. Even though you do not have access to those books at some point in the future, you use what you learned to make decisions today. So you are using the data, but not directly\r\n at the time of the decision.</p>\r\n<p>They want something that has learned about all the data in the past, has some formulas and logic with parameters (perhaps even entered by the user) and constants based on that data, and when presented with TODAY'S data only (eg. today's weather, today's\r\n prior delays, today's number of flights), comes to some prediction about arrival times for planes currently in the air. It will not have access to all the historical data at the time it is used. So it's possible to say &quot;It's raining today, so flights will\r\n be delayed&quot;, but it's not possible to say &quot;It's been raining for the 5th straight day, so flights will be delayed&quot;. You will not know how many days in a row it has been raining, and can only predict that if it's raining today, it's x% more likely to cause\r\n delays today.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18352",
      "postDate": "12/21/2012 15:43:37",
      "content": "<p>Which comes back to my point, essentially: your &quot;x%&quot; IS based on data other than today's observations ...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18356",
      "postDate": "12/21/2012 16:29:12",
      "content": "<p>My understanding was that 14-days of data is provided from which to build a model. For the final, an additional days worth of data is provided with a cutoff time. Data after the cutoff time cannot be used. For all flights remaining in the air after the cutoff\r\n time, predict the runway and gate arrival times (using the 14 days data &#43; additional data from the pre-cutoff time from the extra day).</p>\r\n<p>But, reading the forums I'm sure the above is incorrect!</p>\r\n<p>If the final model is to be based on just one days worth of data prior to the cutoff time then this is not a lot of data to work with.</p>\r\n<p>In a previous note, I mentioned that the challenge is to &quot;... find a scalable algorithm to provide a real-time profile to pilots ...&quot;. Not sure how the above meets the requirements of the challenge.</p>\r\n<p>Realistically the historical flight data could/should stretch to cover at least one season (ie. 3 months ), combined with data from the previous season; with the flight data constantly updated in real-time; and predictions available in real-time.</p>\r\n<p>Is a dynamic problem being asked to be solved as if it were a static problem. Anyway, lots of confusion.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18358",
      "postDate": "12/21/2012 17:37:22",
      "content": "<p>[quote=Konrad Banachewicz;18352]</p>\r\n<p>Which comes back to my point, essentially: your &quot;x%&quot; IS based on data other than today's observations ...</p>\r\n<p>[/quote]</p>\r\n<p>Exactly my point. When you are&nbsp;using ANY historical data to optimise your desicions - you are violating the rules. You are proposing to calculate x% based on historical data. You may do this in realtime or upfront but it makes no difference.</p>\r\n<p>If one CAN use this data to build a model then&nbsp;one can simply call ALL HISTORICAL data as parameters of&nbsp;his model.</p>\r\n<p>Buy I agree with jsink that intentions of organisers was to somehow restrict complexity of the models to make them easily deployable. Problem is that you will not be able to distinguish between model parameters and stored hitorical data and that is why my\r\n point is that all historical data prior the cutoff point should be used to make predictions.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18382",
      "postDate": "12/21/2012 21:18:12",
      "content": "<p>My 2 cents:</p>\r\n<p>the meaning of the rules is that you can use any parameter calculated with the 14days data provided but you can't directly reference in your program those files.<br>\r\nFor example, if you assigned a score to every airport you can use this score in the model, eventually update it with the data of the day of the entry, but can't use the &quot;old files&quot; to rebuild this score from 0.</p>\r\n<p>If it's not so and the organizers say you can't use any parameter calculated on previous data, i would agree with Ivan saying it could be impossible to make any model with some sense.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18383",
      "postDate": "12/21/2012 21:23:10",
      "content": "I am still confused about how to use the &quot;final evaluation&quot; test data. Are we allowed to use prior entries from this test set to predict the future entries? Or must we exclude the whole &quot;final evaluation&quot; test set past and future, using only individual\r\n entries to make separate predictions? Either way, I dont trust they can audit the code &#43; models to ensure the submission does not make use of future data.",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18431",
      "postDate": "12/22/2012 17:56:47",
      "content": "<p>From the admin's responses, here's what I believe the rules are:</p>\r\n<p>1) Models are constructed using <em>all</em> data available before the final test set.</p>\r\n<p>2) In order to make a prediction for a day in the final test set, you are allowed to use your model &#43; only the data for that day.</p>\r\n<p>Admins, is this correct? &nbsp;</p>\r\n<p>If so, this makes the leaderboard a bit inaccurate; I, for one, have been using all the data (including those from test set days) to create my models. &nbsp;Also, if this restriction is correct, the final models will not be as good as they could be. &nbsp;I have noticed\r\n some trends in the data which hold over time spans of ~ 1 day. &nbsp;If we are not allowed to use data directly around the test set day, these trends will go unnoticed ...</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18441",
      "postDate": "12/22/2012 22:18:39",
      "content": "<p>Hi Lonely,</p>\r\n<p>The way I interpret the rules is that we are allowed to use the training set to generate model parameters and only a single day's data to generate estimates. In your case, I don't it's possible, within to the rules, to track multi-day trends extending into\r\n the test set period. That said, I think it's a pretty neat idea.&nbsp;</p>\r\n<p>My assumption is that our models should take as input a single day's worth of data and nothing else. However, model parameters can be derived from the training data.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18447",
      "postDate": "12/23/2012 00:01:45",
      "content": "<p>[quote=ricardo sachez;18441]</p>\r\n<p>My assumption is that our models should take as input a single day's worth of data and nothing else. However, model parameters can be derived from the training data.&nbsp;</p>\r\n<p>[/quote]</p>\r\n<p>Using data &quot;as input&quot; to a model is the same as using data to derive model parameters, is it not?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18618",
      "postDate": "12/26/2012 07:11:04",
      "content": "<p>[quote=Lonely at the Top;18431]</p>\r\n<p>From the admin's responses, here's what I believe the rules are:</p>\r\n<p>1) Models are constructed using <em>all</em> data available before the final test set.</p>\r\n<p>2) In order to make a prediction for a day in the final test set, you are allowed to use your model &#43; only the data for that day.</p>\r\n<p>Admins, is this correct? &nbsp;</p>\r\n<p>If so, this makes the leaderboard a bit inaccurate; I, for one, have been using all the data (including those from test set days) to create my models. &nbsp;Also, if this restriction is correct, the final models will not be as good as they could be. &nbsp;I have noticed\r\n some trends in the data which hold over time spans of ~ 1 day. &nbsp;If we are not allowed to use data directly around the test set day, these trends will go unnoticed ...</p>\r\n<p>[/quote]This is correct. Your model will need to make predictions for each day in the Final Evaluation Set independently.</p>\r\n<p>As you said, this means that results on the Public Leaderboard aren't precisely analogous to results on the Final Evaluation Set. However, this isn't the only way in which the results aren't analogous: you'll be able to use future information from December-January\r\n to make predictions on the Public Leaderboard Set as well, and we can't be certain that people aren't violating the rules in the predictions they are submitting on the public leaderboard set by using external data (such as the actual arrival times for flights\r\n in the test set).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18625",
      "postDate": "12/26/2012 11:31:59",
      "content": "<p>[quote=Ben Hamner;18618]</p>\r\n<p>This is correct. Your model will need to make predictions for each day in the Final Evaluation Set independently.</p>\r\n<p>As you said, this means that results on the Public Leaderboard aren't precisely analogous to results on the Final Evaluation Set. However, this isn't the only way in which the results aren't analogous: you'll be able to use future information from December-January\r\n to make predictions on the Public Leaderboard Set as well, and we can't be certain that people aren't violating the rules in the predictions they are submitting on the public leaderboard set by using external data (such as the actual arrival times for flights\r\n in the test set).</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>Ben, just to make things clearer to me.</p>\r\n<p>Let's suppose I found out from the test dataset that (just as example) for every departure airport there's a coefficent X useful to determine the arrival time for a flight.</p>\r\n<p>So my model says that arrival time (At) depends from this coefficient X, different for every airport, and the scheduled gate departure time (Dt) =&gt; At = X*Dt.</p>\r\n<p>If i understood correctly, i can't use the X that i calculated from the test set but i have to calculate it for each day of the final evaluation set.</p>\r\n<p>In my opinion, this limitation causes more troubles at it solves. For example, let's suppose the cutoff time for a day is 01:00 UTC, for short flights i have very few data to work on, because for example flight plans would be made and changed on the previous\r\n day and so on. If i could use my table of parameters calculated with all the previous data (and, when submitting the model, i'll submit also the way it is created so you can check i didn't use external or future data), i could obtain a much better result.</p>\r\n<p>Thank you&nbsp;</p>\r\n<p>Pierluigi</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18632",
      "postDate": "12/26/2012 16:06:47",
      "content": "<p>Pierluigi, ben Hammer has confirmed that</p>\r\n<p>1) Models are constructed using&nbsp;<em>all</em>&nbsp;data available before the final test set.</p>\r\n<p>2) In order to make a prediction for a day in the final test set, you are allowed to use your model &#43; only the data for that day.</p>\r\n<p>So you can construct your linear model using&nbsp;all data available before the Final Evaluation Set, plus for each day of the Final Evaluation set any data available for this day specifically.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18780",
      "postDate": "12/29/2012 17:24:45",
      "content": "<p>This is a no-brainer.<br>\r\nYou cannot train any model on the same day. As you see, the actual<em>gate</em>arrival and runway_arrival are both hidden. How do you expect to build a model for the same?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18782",
      "postDate": "12/29/2012 19:14:25",
      "content": "<p>Black Magic, you're right this information will be hidden for the flights we will have to predict. However, you may wish to notice this is not hidden for the flights that already took off at cutoff time. This may be usefull.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 17284,
      "author_name": "mcstar",
      "author_url": "",
      "post_date": "11/30/2012 19:36:42",
      "content": "<p>Additionally, it appears that this wording even implies that the model should &quot;start clean&quot; on each test. IOW, it would be an implicit violation if your model has ever been trained on data from the future from the standpoint of the randomly chosen test day.\r\n</p>\r\n<p>This is a condition that may prove nearly impossible to implement since it means you must retrain your model fresh for each test. Any learned information in the model would need to explicitly come from past days only. Since it's impossible to identify the\r\n source of learned data in most models, we must either re-train on every test, or assume that the test period will be chosen from a hidden set of dates we won't have access to in the training set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17296,
      "author_name": "dchudz",
      "author_url": "",
      "post_date": "11/30/2012 21:35:04",
      "content": "<p>Our apologies -- the wiki and data page were wrong; the submission instructions page was correct. Now they are all correct.</p>\r\n<p>Your model must be structured to make final test data set predictions based on no days in the final test data set other than the appropriate one (none prior or later are allowed).\r\n</p>\r\n<p>This rule only applies to the final test data set. Note that some of the training data will be from a period later than the public leaderboard test data set. It is fine to use this data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17317,
      "author_name": "intaka",
      "author_url": "",
      "post_date": "12/01/2012 11:31:33",
      "content": "<p>&quot;For each day in the test period (first for the public leaderboard, and later for the final evaluation), we will select a random time (uniformly chosen between 9am EST and 9pm EST) and select all of the flights in the air at that cutoff time. You will be\r\n provided with relevant data for each day that would be available at the chosen cutoff time.&quot;</p>\r\n<p><strong>&quot;Your model must be structured so that it makes each test day's final test data set predictions based on no information in the final evaluation test data other than the information from that day, which will be in an appropriately named folder.&nbsp;</strong>(Reworded\r\n for clarification on 11/30/2012. See&nbsp;<a href=\"https://www.gequest.com/c/flight/forums/t/3220/contradiction-in-rules-for-final-model\">forum</a>&nbsp;for explanation.&quot;</p>\r\n<p>It would be useful to people who are new to the flight data domain if the above could be explained with clarity by the organizers. Walk it through with an example or something that assists understanding. &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17334,
      "author_name": "joexjmmvhm",
      "author_url": "",
      "post_date": "12/02/2012 05:33:30",
      "content": "<ol>\r\n<li>I am guessing that for each test day we can use all data (including Arrival Times Scheduled and Actual) for that day up until a random cutoff time and then we will predict the arrival times for the rest of that day of only the aircraft that are in the air\r\n at the time of the cutoff. This would mean no extra points for good predictions of how long it takes the plane to board, taxi down the runway, and get into the air.\r\n</li></ol>\r\n<p>Can anyone tell me if this sounds correct?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17376,
      "author_name": "marka25933",
      "author_url": "",
      "post_date": "12/04/2012 05:53:19",
      "content": "<p>Just how far does this use of time constrained data go? Given that, almost by definition, any model used will have been derived from data outside this period, including constants, structure, formulae etc.&nbsp;</p>\r\n<p>Some further clarity would be helpful to keep pedantry under control. An example could well help too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17383,
      "author_name": "farstars",
      "author_url": "",
      "post_date": "12/04/2012 17:32:58",
      "content": "<p>I'm still little bit confused about the objective of the competition.</p>\r\n<p>Let's assume I am a pilot flying, right now, towards my destination airport.</p>\r\n<p>A wanted model, such as a final model (whose prediction performance should have been measured at a particular cut-off time), will give me a fairly good estimation for my actual arrival time on that airport. So, I have now better estimation for my arrival\r\n time which differs from the original flight plan.</p>\r\n<p>Then, how can I use this knowledge to improve 'efficiency' of my flight? Do I have to increase/reduce my plane's speed to stick closer to the original plan?</p>\r\n<p>Could anyone, please, clarify the meaning of the objective to predict arrival time in terms of &quot;Make flying more efficient?&quot;</p>\r\n<p>Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17418,
      "author_name": "dchudz",
      "author_url": "",
      "post_date": "12/05/2012 19:00:46",
      "content": "<p>Joexjmmvhm-- that is correct.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17419,
      "author_name": "dchudz",
      "author_url": "",
      "post_date": "12/05/2012 19:04:23",
      "content": "<p>MarkA-- your models (parameters,etc.) will be based on the training data. This model then needs to be able to look at one day of the final test data (in isolation) and make predictions. So the sequence is:</p>\r\n<p>1) Build models based on training data, submitting to public leaderboard<br>\r\n2) Finalize code based on full training data set -- submit code -- code (which will include all model parameters, etc.) must be able to look at one day of test data in isolation and make predictions for that day<br>\r\n3) Final test data is released<br>\r\n4) Submit predictions for test data, using submitted code<br>\r\n5) Winners' code will be verified for compliance with these rules</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17420,
      "author_name": "dchudz",
      "author_url": "",
      "post_date": "12/05/2012 19:06:30",
      "content": "<p>Farstar-- the pilot wouldn't use these predictions directly. Instead, these predictions will feed into a broader optimization of parameters like the ones you mentioned, like speed. You can't optimize over the outcomes until you know which situations will\r\n lead to which outcomes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18332,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "12/21/2012 12:34:40",
      "content": "<p>David, you wrote:</p>\r\n<blockquote>\r\n<p>1) Build models based on training data, submitting to public leaderboard<br>\r\n2) Finalize code based on full training data set -- submit code -- code (which will include all model parameters, etc.) must be able to look at one day of test data in isolation and make predictions for that day</p>\r\n</blockquote>\r\n<p>You see, when you &quot;train&quot; a model, it can actually memorise all the data it has seen. So your statement that you &quot;allow to train model on all data and do not allow to use all data before cutoff date for scoring&quot; does not make much sense.</p>\r\n<p>Let me give you an example - you are building a KNN model that actually memorises all data in it and then simply finds closest neighbours. Although you might think that you are using only current record (data from the current day) for scoring, other records\r\n from previous days are used as well because you need to calculate distances to them.</p>\r\n<p>Because you allow to train on all data, my model can memorise it all as a parameters that you allow to store.</p>\r\n<p>Let me give you another example. For this challenge I want to build this type of naive model:</p>\r\n<p>1. For each airport I canculate average delay of arrival (average difference between actual and estimated arrrival times).</p>\r\n<p>2. For each flight in the air I simply subtract estimated arrival time from the cutoff time and add average delay for the destination airport.</p>\r\n<p>So here is the question. You say that this naive model is not allowed, because to calculate it I am using average delays for airports that I calculated using all the data.To my mind this rule does not make practical sence because it contradicts with your\r\n statement that you allow to save model parameters.</p>\r\n<p>Because any parameters are stored as bytes and all the data from previous days is stored as bytes you cannot allow one thing and prohibit another.</p>\r\n<p>Do you agree?</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18333,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "12/21/2012 12:59:17",
      "content": "<p>Hear, hear... I was kinda wondering about the same thing (i.e. ban on using data from other days vs permission to use parameters trained on those days)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18335,
      "author_name": "bjg9846",
      "author_url": "",
      "post_date": "12/21/2012 13:25:57",
      "content": "<p>I have been interpreting it this way. &nbsp;Any data that is available from the competition before the model submission deadline (and prior to the final evaluation set is released) is fair game for the models. &nbsp;Then, when the final data set is released, the model\r\n execution should be restricted to the one day of test data. &nbsp;So, a KNN model could contain all the training data, but should not add any of the instances from the other days in the final evaluation set to the neighborhood.</p>\r\n<p>For example, a model should produce the exact same predictions for day 4 in the final evaluation set, regardless of how many other days of data are provided in that final evaluation set.</p>\r\n<p>Is this the intent?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18339,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "12/21/2012 13:56:12",
      "content": "<p>Fine: let's take a linear regression model as an example. If I have estimated the parameters on e.g. all days from november and then I try to use this model taking only december 4th as input - I am still - implicitly - using the other observations, because\r\n that affects what my regression parameters look like - so technically, I am in violation of the rules. If you want to avoid this, then training data should not be used at all...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18342,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "12/21/2012 14:18:05",
      "content": "<p>[quote=BJG_;18335]</p>\r\n<p>I have been interpreting it this way. &nbsp;Any data that is available from the competition before the model submission deadline (and prior to the final evaluation set is released) is fair game for the models. &nbsp;Then, when the final data set is released, the model\r\n execution should be restricted to the one day of test data. &nbsp;So, a KNN model could contain all the training data, but should not add any of the instances from the other days in the final evaluation set to the neighborhood.</p>\r\n<p>For example, a model should produce the exact same predictions for day 4 in the final evaluation set, regardless of how many other days of data are provided in that final evaluation set.</p>\r\n<p>Is this the intent?</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>This intent makes no sense. You can imagine a model as a black box loaded with some universal knowledge that makes predictions. You come to this black box, open it, put your new day data into it, step back and then ask &quot;Dear box, when do you think this plane\r\n that I have just told you about, will land?&quot;. The box uses ALL of its knowledge to answer this question. Both your current data and everything it has seen before.</p>\r\n<p>Here is an example - IBM Watson machine that has beaten humans in Jeopardy game. It was so clever because it had access to HUGE amounts of text data and it could search through it in real time and find answers.</p>\r\n<p>Another example - you are playing chess with AI. You are loosing because it has millions of strategies loaded in it and it is optimising its chanses to win by predicting which strategy is the best. And you say: &quot;Stop doing that! This is not fair! Please,\r\n do not use all your strategies that you have learned from previous games. You can use data only from this game!&quot; =)&nbsp; Ha-ha, that would be funny =)</p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18344,
      "author_name": "bjg9846",
      "author_url": "",
      "post_date": "12/21/2012 14:28:51",
      "content": "<p>I'm guessing the primary intent is to avoid using the future data in the prediction. &nbsp;They do not want a model to use data from Day 14 in the final evaluation set to make predictions for Day 4.&nbsp;</p>\r\n<p>There is also a chance that the customer is risk adverse in this area and does not want to have a model whose parameters change automatically after it is deployed. &nbsp;They may plan to go through some testing process prior to deploying the model and want any\r\n changes to the model to go through a similar testing process (not learned automatically). &nbsp;</p>\r\n<p>&nbsp;I am just conjecturing though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18350,
      "author_name": "jsink71991",
      "author_url": "",
      "post_date": "12/21/2012 15:28:17",
      "content": "<p>It's not like Jeopardy, but it may be like the chess example.<br>\r\nI think the thing to do is to look at how this will be used in the real world. <br>\r\nThey do not want the tool to go through all of history and also include today's data and come to some analysis.<br>\r\nIt will not have access to all this data, nor will it have all this time to do analysis and arrive at an answer.</p>\r\n<p>You go to school for x number of years, and you read many books and learn from them. Even though you do not have access to those books at some point in the future, you use what you learned to make decisions today. So you are using the data, but not directly\r\n at the time of the decision.</p>\r\n<p>They want something that has learned about all the data in the past, has some formulas and logic with parameters (perhaps even entered by the user) and constants based on that data, and when presented with TODAY'S data only (eg. today's weather, today's\r\n prior delays, today's number of flights), comes to some prediction about arrival times for planes currently in the air. It will not have access to all the historical data at the time it is used. So it's possible to say &quot;It's raining today, so flights will\r\n be delayed&quot;, but it's not possible to say &quot;It's been raining for the 5th straight day, so flights will be delayed&quot;. You will not know how many days in a row it has been raining, and can only predict that if it's raining today, it's x% more likely to cause\r\n delays today.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18352,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "12/21/2012 15:43:37",
      "content": "<p>Which comes back to my point, essentially: your &quot;x%&quot; IS based on data other than today's observations ...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18356,
      "author_name": "intaka",
      "author_url": "",
      "post_date": "12/21/2012 16:29:12",
      "content": "<p>My understanding was that 14-days of data is provided from which to build a model. For the final, an additional days worth of data is provided with a cutoff time. Data after the cutoff time cannot be used. For all flights remaining in the air after the cutoff\r\n time, predict the runway and gate arrival times (using the 14 days data &#43; additional data from the pre-cutoff time from the extra day).</p>\r\n<p>But, reading the forums I'm sure the above is incorrect!</p>\r\n<p>If the final model is to be based on just one days worth of data prior to the cutoff time then this is not a lot of data to work with.</p>\r\n<p>In a previous note, I mentioned that the challenge is to &quot;... find a scalable algorithm to provide a real-time profile to pilots ...&quot;. Not sure how the above meets the requirements of the challenge.</p>\r\n<p>Realistically the historical flight data could/should stretch to cover at least one season (ie. 3 months ), combined with data from the previous season; with the flight data constantly updated in real-time; and predictions available in real-time.</p>\r\n<p>Is a dynamic problem being asked to be solved as if it were a static problem. Anyway, lots of confusion.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18358,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "12/21/2012 17:37:22",
      "content": "<p>[quote=Konrad Banachewicz;18352]</p>\r\n<p>Which comes back to my point, essentially: your &quot;x%&quot; IS based on data other than today's observations ...</p>\r\n<p>[/quote]</p>\r\n<p>Exactly my point. When you are&nbsp;using ANY historical data to optimise your desicions - you are violating the rules. You are proposing to calculate x% based on historical data. You may do this in realtime or upfront but it makes no difference.</p>\r\n<p>If one CAN use this data to build a model then&nbsp;one can simply call ALL HISTORICAL data as parameters of&nbsp;his model.</p>\r\n<p>Buy I agree with jsink that intentions of organisers was to somehow restrict complexity of the models to make them easily deployable. Problem is that you will not be able to distinguish between model parameters and stored hitorical data and that is why my\r\n point is that all historical data prior the cutoff point should be used to make predictions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18382,
      "author_name": "pierluigivinciguerra",
      "author_url": "",
      "post_date": "12/21/2012 21:18:12",
      "content": "<p>My 2 cents:</p>\r\n<p>the meaning of the rules is that you can use any parameter calculated with the 14days data provided but you can't directly reference in your program those files.<br>\r\nFor example, if you assigned a score to every airport you can use this score in the model, eventually update it with the data of the day of the entry, but can't use the &quot;old files&quot; to rebuild this score from 0.</p>\r\n<p>If it's not so and the organizers say you can't use any parameter calculated on previous data, i would agree with Ivan saying it could be impossible to make any model with some sense.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18383,
      "author_name": "studentcjd",
      "author_url": "",
      "post_date": "12/21/2012 21:23:10",
      "content": "I am still confused about how to use the &quot;final evaluation&quot; test data. Are we allowed to use prior entries from this test set to predict the future entries? Or must we exclude the whole &quot;final evaluation&quot; test set past and future, using only individual\r\n entries to make separate predictions? Either way, I dont trust they can audit the code &#43; models to ensure the submission does not make use of future data.",
      "votes": null,
      "replies": []
    },
    {
      "id": 18431,
      "author_name": "lonelyatthetop",
      "author_url": "",
      "post_date": "12/22/2012 17:56:47",
      "content": "<p>From the admin's responses, here's what I believe the rules are:</p>\r\n<p>1) Models are constructed using <em>all</em> data available before the final test set.</p>\r\n<p>2) In order to make a prediction for a day in the final test set, you are allowed to use your model &#43; only the data for that day.</p>\r\n<p>Admins, is this correct? &nbsp;</p>\r\n<p>If so, this makes the leaderboard a bit inaccurate; I, for one, have been using all the data (including those from test set days) to create my models. &nbsp;Also, if this restriction is correct, the final models will not be as good as they could be. &nbsp;I have noticed\r\n some trends in the data which hold over time spans of ~ 1 day. &nbsp;If we are not allowed to use data directly around the test set day, these trends will go unnoticed ...</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18441,
      "author_name": "rickyars",
      "author_url": "",
      "post_date": "12/22/2012 22:18:39",
      "content": "<p>Hi Lonely,</p>\r\n<p>The way I interpret the rules is that we are allowed to use the training set to generate model parameters and only a single day's data to generate estimates. In your case, I don't it's possible, within to the rules, to track multi-day trends extending into\r\n the test set period. That said, I think it's a pretty neat idea.&nbsp;</p>\r\n<p>My assumption is that our models should take as input a single day's worth of data and nothing else. However, model parameters can be derived from the training data.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18447,
      "author_name": "studentcjd",
      "author_url": "",
      "post_date": "12/23/2012 00:01:45",
      "content": "<p>[quote=ricardo sachez;18441]</p>\r\n<p>My assumption is that our models should take as input a single day's worth of data and nothing else. However, model parameters can be derived from the training data.&nbsp;</p>\r\n<p>[/quote]</p>\r\n<p>Using data &quot;as input&quot; to a model is the same as using data to derive model parameters, is it not?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18618,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "12/26/2012 07:11:04",
      "content": "<p>[quote=Lonely at the Top;18431]</p>\r\n<p>From the admin's responses, here's what I believe the rules are:</p>\r\n<p>1) Models are constructed using <em>all</em> data available before the final test set.</p>\r\n<p>2) In order to make a prediction for a day in the final test set, you are allowed to use your model &#43; only the data for that day.</p>\r\n<p>Admins, is this correct? &nbsp;</p>\r\n<p>If so, this makes the leaderboard a bit inaccurate; I, for one, have been using all the data (including those from test set days) to create my models. &nbsp;Also, if this restriction is correct, the final models will not be as good as they could be. &nbsp;I have noticed\r\n some trends in the data which hold over time spans of ~ 1 day. &nbsp;If we are not allowed to use data directly around the test set day, these trends will go unnoticed ...</p>\r\n<p>[/quote]This is correct. Your model will need to make predictions for each day in the Final Evaluation Set independently.</p>\r\n<p>As you said, this means that results on the Public Leaderboard aren't precisely analogous to results on the Final Evaluation Set. However, this isn't the only way in which the results aren't analogous: you'll be able to use future information from December-January\r\n to make predictions on the Public Leaderboard Set as well, and we can't be certain that people aren't violating the rules in the predictions they are submitting on the public leaderboard set by using external data (such as the actual arrival times for flights\r\n in the test set).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18625,
      "author_name": "pierluigivinciguerra",
      "author_url": "",
      "post_date": "12/26/2012 11:31:59",
      "content": "<p>[quote=Ben Hamner;18618]</p>\r\n<p>This is correct. Your model will need to make predictions for each day in the Final Evaluation Set independently.</p>\r\n<p>As you said, this means that results on the Public Leaderboard aren't precisely analogous to results on the Final Evaluation Set. However, this isn't the only way in which the results aren't analogous: you'll be able to use future information from December-January\r\n to make predictions on the Public Leaderboard Set as well, and we can't be certain that people aren't violating the rules in the predictions they are submitting on the public leaderboard set by using external data (such as the actual arrival times for flights\r\n in the test set).</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>Ben, just to make things clearer to me.</p>\r\n<p>Let's suppose I found out from the test dataset that (just as example) for every departure airport there's a coefficent X useful to determine the arrival time for a flight.</p>\r\n<p>So my model says that arrival time (At) depends from this coefficient X, different for every airport, and the scheduled gate departure time (Dt) =&gt; At = X*Dt.</p>\r\n<p>If i understood correctly, i can't use the X that i calculated from the test set but i have to calculate it for each day of the final evaluation set.</p>\r\n<p>In my opinion, this limitation causes more troubles at it solves. For example, let's suppose the cutoff time for a day is 01:00 UTC, for short flights i have very few data to work on, because for example flight plans would be made and changed on the previous\r\n day and so on. If i could use my table of parameters calculated with all the previous data (and, when submitting the model, i'll submit also the way it is created so you can check i didn't use external or future data), i could obtain a much better result.</p>\r\n<p>Thank you&nbsp;</p>\r\n<p>Pierluigi</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18632,
      "author_name": "jay670790",
      "author_url": "",
      "post_date": "12/26/2012 16:06:47",
      "content": "<p>Pierluigi, ben Hammer has confirmed that</p>\r\n<p>1) Models are constructed using&nbsp;<em>all</em>&nbsp;data available before the final test set.</p>\r\n<p>2) In order to make a prediction for a day in the final test set, you are allowed to use your model &#43; only the data for that day.</p>\r\n<p>So you can construct your linear model using&nbsp;all data available before the Final Evaluation Set, plus for each day of the Final Evaluation set any data available for this day specifically.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18780,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "12/29/2012 17:24:45",
      "content": "<p>This is a no-brainer.<br>\r\nYou cannot train any model on the same day. As you see, the actual<em>gate</em>arrival and runway_arrival are both hidden. How do you expect to build a model for the same?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18782,
      "author_name": "jay670790",
      "author_url": "",
      "post_date": "12/29/2012 19:14:25",
      "content": "<p>Black Magic, you're right this information will be hidden for the flights we will have to predict. However, you may wish to notice this is not hidden for the flights that already took off at cutoff time. This may be usefull.&nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "17281": "",
    "17284": "",
    "17296": "",
    "17317": "",
    "17334": "",
    "17376": "",
    "17383": "",
    "17418": "",
    "17419": "",
    "17420": "",
    "18332": "",
    "18333": "",
    "18335": "",
    "18339": "",
    "18342": "",
    "18344": "",
    "18350": "",
    "18352": "",
    "18356": "",
    "18358": "",
    "18382": "",
    "18383": "",
    "18431": "",
    "18441": "",
    "18447": "",
    "18618": "",
    "18625": "",
    "18632": "",
    "18780": "",
    "18782": ""
  },
  "source": "meta"
}