{
  "id": 4183,
  "title": "Congratulations!",
  "url": "/competitions/flight/discussion/4183",
  "author_name": "",
  "post_date": "2013-04-03T12:53:15.993Z",
  "votes": null,
  "comment_count": 23,
  "views": 21586,
  "content": "<p>Congratulations to the winners of Flight Quest Phase 1!</p>\r\n<p>Nice Job!</p>\r\n<p>Team Gxav &amp; * getting 4.2 for gate and 3.2 for runway arrival!</p>\r\n<p>Good work!</p>\r\n<p>Jules</p>",
  "messages": [
    {
      "id": "22058",
      "postDate": "04/03/2013 12:53:15",
      "content": "<p>Congratulations to the winners of Flight Quest Phase 1!</p>\r\n<p>Nice Job!</p>\r\n<p>Team Gxav &amp; * getting 4.2 for gate and 3.2 for runway arrival!</p>\r\n<p>Good work!</p>\r\n<p>Jules</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22060",
      "postDate": "04/03/2013 13:00:12",
      "content": "<p>Congratulations!&nbsp;<span style=\"line-height:1.4em\">Will the other teams results be published?</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22061",
      "postDate": "04/03/2013 13:01:54",
      "content": "<p>Congratuations to all winners! Would be curious to see the others results as well!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22062",
      "postDate": "04/03/2013 13:03:19",
      "content": "<p>Congrats everyone!&nbsp; It was a fun competition.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22064",
      "postDate": "04/03/2013 13:30:34",
      "content": "<p>Congratulations to the winners!&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>@BreakfastPirate:&nbsp;I would say rather &quot;hard competition&quot; than fun competition :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22067",
      "postDate": "04/03/2013 14:03:47",
      "content": "<p>@beluga &amp; @BreakfastPirate: This was both &quot;fun&quot; and &quot;hard competition&quot; :).</p>\r\n<p>@Peter: I watched your fast progress with awe. Given your late start your results were incredible.</p>\r\n<p>Congrats to all who managed to work with this data. It is also exciting to see that how different teams used totally different strategies.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22071",
      "postDate": "04/03/2013 14:22:14",
      "content": "<p>Now everytime I see an airplane fly overhead, I find myself thinking, &quot;I wonder what that airplane's ERA is?&nbsp; I wonder how recently its ERA was updated?&nbsp; I wonder if I took current time &#43; (distance&nbsp;to airport / airspeed) if that would be more than ERA or\r\n less than ERA?&nbsp; I wonder if it has&nbsp;been diverted ... in which case I don't need to worry about it?&quot;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22072",
      "postDate": "04/03/2013 14:30:27",
      "content": "<p>@BreakfastPirate: same here - especially about diverted flights ;-)</p>\r\n<p>@Pawel: Thanks - I did a poor job on model selection, though - I mostly used the LB for that since I had a hard time doing internal validations that correlated with the LB - it seems I learned the lesson the hard way</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22079",
      "postDate": "04/03/2013 16:06:25",
      "content": "<p><span>Congratulations to the winners!&nbsp;</span></p>\r\n<p><span>@Peter: Thank you for your implementation of gradient tree boosting!</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22083",
      "postDate": "04/03/2013 16:19:23",
      "content": "<p>Peter:<br>\r\nWhat is LB?</p>\r\n<p><span style=\"line-height:1.4em\">What kinds of features did you use? I found that using raw averages in GBM overfits data terribly? Did you follow any trick for the variables?&nbsp;</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22088",
      "postDate": "04/03/2013 18:14:44",
      "content": "<p>@Black Magic: LB stands for Leaderboard.</p>\r\n<p>I am also quite curious about more details on approaches people took in this competition, and especially how they estimate the influence of the various data sources, such as ASDI. Since not having found any time from short after starting with this competition,\r\n I ended up with a model only using data from both &quot;flight history&quot; files - which is why I am suprised about finished this close to the winners.</p>\r\n<p>What I did was to implement various heuristics to handle the rather sparse and often outdated updates from the events files for finding bounds and predictions on the runways arrival times. Then, using some basic features such as expected arrival and airline/airport\r\n statistics, I estimated the difference between runway and gate arrival.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22090",
      "postDate": "04/03/2013 18:29:43",
      "content": "<p>@stefan:</p>\r\n<p>The data itself has a big&nbsp;variance. For the best benchmark (which can be regarded as a feature in this competition), RMSEs are signifcantly different on different sets:</p>\r\n<p>set &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;RMSE of benchmark</p>\r\n<p><span>InitialTrainingSet_rev1</span>&nbsp; &nbsp; &nbsp; &nbsp;9.7</p>\r\n<p><span>AugmentedTrainingSet2 &nbsp; &nbsp;around 11</span></p>\r\n<p><span>FinalEvaluationSet &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 13</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22099",
      "postDate": "04/03/2013 19:07:09",
      "content": "<p>I mainly relied on: 1) most recent ERA from flight history 2) most recent ERA from ASDI flight plan 3) estimated flight time remaining based on most recent lat/lon from ASDI position. I then spent a LOT of time trying to determine heuristics for when one\r\n or more of these were unreliable/erroneous. I threw in a couple variables at the end that measured whether the flight was<br>\r\nprimarily going North, South, East, or West. My final model for runway arrival (combo of SVM and Random Forest) only used 12 variables.&nbsp; Then for gate arrival I used a 3-variable linear regression model to determine runway-to-gate difference.</p>\r\n<p>So it ended up being heavy on data-cleaning/error-detection and light on modeling.</p>\r\n<p>Steve</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22101",
      "postDate": "04/03/2013 19:26:33",
      "content": "<p>@BlackMagic: I used features from all files except the weather forecasts (only current weather) and the flightplans (I did use the last waypoint before the arrival airport).</p>\r\n<p>My overal approach was to use a custom gradient boosting machine that was able to predict multiple outputs; thus, I didn't train individual models for runway and gate arrival or stacked the gate model on top of the runway model. Furthermore, I didn't do\r\n any data cleansing and post-processing of my predicitons. I used the expected arrival benchmark as my initial model, thus my gbm started with the EAB predictions and made&nbsp;</p>\r\n<p>The most important features are derived from the various times (arrival, departure, scheduled, block). Then comes longitude/latitude information of the arrival airport - I think this is used by my model to group arrival airports to certain regions (and to\r\n pool information rather than using individual arrival airports for which data can be quite sparse).</p>\r\n<p>Another pretty strong feature was the minimum flight time given the last known location of the airplane and its cruising speed. I experimented with the approach angle as well but it didn't seem to help much - maybe because the model already uses the ASDI\r\n location as a surogat. I created numerous additional features by computing time differences (e.g. minimum flight time - time to estimated runway arrival).</p>\r\n<p>According to my model, weather information was only marginally useful; the model didn't pick up any signals in ground delay and deicing reports (too sparse data?).</p>\r\n<p>You can find the feature importance plot of the top 50 most important features (relative scores) in the attachment.</p>\r\n<p>I also included some partial dependence plots - I think they are of no use anymore :-/</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22111",
      "postDate": "04/03/2013 23:06:17",
      "content": "<p>Like all the other posters, congratulations to all the winners.&nbsp; Inspiration for Phase 2.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22147",
      "postDate": "04/04/2013 09:45:25",
      "content": "<p>The methods of the winners do not seem all that different to me.&nbsp;<span style=\"line-height:1.4em\">Here is mine:</span></p>\r\n<p><span style=\"line-height:1.4em\">Construct a smoothed, gridded speed map indexed by lat/lon/direction from asdi data giving more weight to recent observations. This is much like a gaussian random field. To predict runway arrival, simulate crossing map by\r\n the route in the latest plan. Fall back on ERA, etc if there is no positions or plan data.</span></p>\r\n<p><span style=\"line-height:1.4em\">The speed map seems to pick up the weather factors nicely. Mostly the wind, I guess.</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22148",
      "postDate": "04/04/2013 09:51:29",
      "content": "<p>@Gabor This sounds quite interesting indeed - can you elaborate more?</p>\r\n<p>It also seems to me that most people made predictions in isolation (the prediction of flight A has not influence on the prediciton of flight B) - has anybody incorporated such inter-dependencies?</p>\r\n<p>I added the scheduled arrival rate at certain intervals (30min, 15min) otherwise predictions are totally independent.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22150",
      "postDate": "04/04/2013 10:18:42",
      "content": "<p>@Peter: You are right. The flights A and B can be interdependent if they their predicted arrival times coincide in time. To know if they coincide you have to make some initial predictions. So there is a loop hole. In theory you could do an infinite recursion\r\n where the predictions interact with each other: estimations -&gt; predictions -&gt; estimations -&gt; etc.... This is closer to a simulation. Maybe our implementation of this logic was not right but I didn't notice much improvement using interdependent variables.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22153",
      "postDate": "04/04/2013 12:33:25",
      "content": "<p>@Gabor I used an approach very similar to what you described. I used lat/lon/altitude from asdi data giving more weight to recent observations. This worked well for flights near the destination. When falling back to ERA, I applied a correction to the estimate,\r\n depending on runway approach direction, and this adjustment made a good improvement to the score.</p>\r\n<p>@Peter I have some code that makes predictions based on other flights' estimates on the same approach airway. If I have a good estimate for a flight, I'd predict the other flights following it by adding the time to cover the distance between their positions.\r\n This worked well for cases when flights go in circles waiting for the queue to clear.</p>\r\n<p>@Pawel If my model outputs the same time for several flights, it'd add a delay to form up a valid queue with some intervals (typically 1 minute). This gave me only a slight improvement to the score. I haven't used a loop, just doing this once.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22164",
      "postDate": "04/04/2013 16:22:43",
      "content": "<p>Congratulations to the winners.</p>\r\n<p>I mostly used obvious time features such as difference between expected and scheduled, and last update from actual departure and mean delay of the airports as well as expected arrival (ERA, EGA and estimatedarrival from asdiflightplan). As for asdiposition,\r\n I used only altitude and groundspeed. Based on this thread, I'd better use lat/long. My final model was made by gbm. I used predicted runway arrival as an input to prediction of gate arrival.</p>\r\n<p>Since I started writing code from Feb 1st, I couldn't make decent submission at the time of the first deadline, Feb 8th. My internal validation on PLB at the time of Feb 15th only using information available before 2/8 was about 6.3 ~ 6.4. I didn't expect\r\n I would finish in top 10 since I couldn't exceed the hurdle of 6.0. I was largely motivated by the forum thread</p>\r\n<p>https://www.gequest.com/c/flight/forums/t/3712/do-you-think-is-too-late-to</p>\r\n<p>and Pawel was right. I couldn't finish in the money...</p>\r\n<p>&nbsp;</p>\r\n<p>To kaggle admin:</p>\r\n<p>I'm sorry for creating the thread to try to reveal or discuss the result the other day. I should have read the embargo thread.</p>\r\n<p>&nbsp;</p>\r\n<p>And thank you for hosting great and really fun competition, Ben.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22468",
      "postDate": "04/09/2013 05:27:19",
      "content": "<p>Compared with the approaches of the other competitors I'm afraid my approach was quite plain jane. I simply had two independent gradient boosted regressors for the gate and runway predictions and fed them the same features. I spent most of my time trying\r\n to think of creative features, in the end I had about 1900 total features of which only 1100 were used by the estimator. I used virtually all the available data with the exception of the ATSCC data. I didn't spend much time doing feature selection/optimization.\r\n Instead I tried to feed the algorithms as much training data as possible. In the end I wanted to use 310,000 training flights but I could only handle 260,000 in memory at one time so that's what I went with. I didn't spend much time trying to deal with outliers,\r\n I just used huber loss with a very high alpha value.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22483",
      "postDate": "04/09/2013 10:18:11",
      "content": "<p>Like Jacques, we decided to discard our features generated from ATSCC data as we didn't find them useful at the early stage of the competition (too few events).</p>\r\n<p>Our solution is a blend of GBMs and RFs but a simpler solution would have provided us with similar accuracy: 7.987. This solution consists of a GBM to model the ERA error and a RF to model the taxi time (difference between runway and gate arrival times).\r\n Our nb of training examples is similar to Jacques' nb.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22527",
      "postDate": "04/09/2013 21:11:46",
      "content": "<p>I used a relatively small number of features (56) and tried to squeeze out as much as possible from them. I generated a random cutoff time for each flight in order to have more training examples. I trained my final model on 850,000 flights (split into 3\r\n parts).</p>\r\n<p>The most informative features in my model were direct estimates of the arrival times (with some cleaning applied). The second best ones were categorical features associated with the flights (e.g. destination, airline, destination x airline, destination x\r\n origin, etc.). I ended up with implementing my own sparse ridge regression solver, because scikit-learn's did not converge on this data.</p>\r\n<p>I also used weather and positional features. The most useful weather feature was the destination airport's most recent &quot;present conditions&quot; value.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "35948",
      "postDate": "12/07/2013 15:47:33",
      "content": "<p>I want to analyze this data for education process. Can anyone help on how to start the process of cleaning the data and suggest some techniques.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 22060,
      "author_name": "iamivo",
      "author_url": "",
      "post_date": "04/03/2013 13:00:12",
      "content": "<p>Congratulations!&nbsp;<span style=\"line-height:1.4em\">Will the other teams results be published?</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22061,
      "author_name": "pprett",
      "author_url": "",
      "post_date": "04/03/2013 13:01:54",
      "content": "<p>Congratuations to all winners! Would be curious to see the others results as well!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22062,
      "author_name": "breakfastpirate",
      "author_url": "",
      "post_date": "04/03/2013 13:03:19",
      "content": "<p>Congrats everyone!&nbsp; It was a fun competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22064,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "04/03/2013 13:30:34",
      "content": "<p>Congratulations to the winners!&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>@BreakfastPirate:&nbsp;I would say rather &quot;hard competition&quot; than fun competition :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22067,
      "author_name": "paweljankiewicz",
      "author_url": "",
      "post_date": "04/03/2013 14:03:47",
      "content": "<p>@beluga &amp; @BreakfastPirate: This was both &quot;fun&quot; and &quot;hard competition&quot; :).</p>\r\n<p>@Peter: I watched your fast progress with awe. Given your late start your results were incredible.</p>\r\n<p>Congrats to all who managed to work with this data. It is also exciting to see that how different teams used totally different strategies.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22071,
      "author_name": "breakfastpirate",
      "author_url": "",
      "post_date": "04/03/2013 14:22:14",
      "content": "<p>Now everytime I see an airplane fly overhead, I find myself thinking, &quot;I wonder what that airplane's ERA is?&nbsp; I wonder how recently its ERA was updated?&nbsp; I wonder if I took current time &#43; (distance&nbsp;to airport / airspeed) if that would be more than ERA or\r\n less than ERA?&nbsp; I wonder if it has&nbsp;been diverted ... in which case I don't need to worry about it?&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22072,
      "author_name": "pprett",
      "author_url": "",
      "post_date": "04/03/2013 14:30:27",
      "content": "<p>@BreakfastPirate: same here - especially about diverted flights ;-)</p>\r\n<p>@Pawel: Thanks - I did a poor job on model selection, though - I mostly used the LB for that since I had a hard time doing internal validations that correlated with the LB - it seems I learned the lesson the hard way</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22079,
      "author_name": "songgc",
      "author_url": "",
      "post_date": "04/03/2013 16:06:25",
      "content": "<p><span>Congratulations to the winners!&nbsp;</span></p>\r\n<p><span>@Peter: Thank you for your implementation of gradient tree boosting!</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22083,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/03/2013 16:19:23",
      "content": "<p>Peter:<br>\r\nWhat is LB?</p>\r\n<p><span style=\"line-height:1.4em\">What kinds of features did you use? I found that using raw averages in GBM overfits data terribly? Did you follow any trick for the variables?&nbsp;</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22088,
      "author_name": "shenss",
      "author_url": "",
      "post_date": "04/03/2013 18:14:44",
      "content": "<p>@Black Magic: LB stands for Leaderboard.</p>\r\n<p>I am also quite curious about more details on approaches people took in this competition, and especially how they estimate the influence of the various data sources, such as ASDI. Since not having found any time from short after starting with this competition,\r\n I ended up with a model only using data from both &quot;flight history&quot; files - which is why I am suprised about finished this close to the winners.</p>\r\n<p>What I did was to implement various heuristics to handle the rather sparse and often outdated updates from the events files for finding bounds and predictions on the runways arrival times. Then, using some basic features such as expected arrival and airline/airport\r\n statistics, I estimated the difference between runway and gate arrival.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22090,
      "author_name": "songgc",
      "author_url": "",
      "post_date": "04/03/2013 18:29:43",
      "content": "<p>@stefan:</p>\r\n<p>The data itself has a big&nbsp;variance. For the best benchmark (which can be regarded as a feature in this competition), RMSEs are signifcantly different on different sets:</p>\r\n<p>set &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;RMSE of benchmark</p>\r\n<p><span>InitialTrainingSet_rev1</span>&nbsp; &nbsp; &nbsp; &nbsp;9.7</p>\r\n<p><span>AugmentedTrainingSet2 &nbsp; &nbsp;around 11</span></p>\r\n<p><span>FinalEvaluationSet &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 13</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22099,
      "author_name": "breakfastpirate",
      "author_url": "",
      "post_date": "04/03/2013 19:07:09",
      "content": "<p>I mainly relied on: 1) most recent ERA from flight history 2) most recent ERA from ASDI flight plan 3) estimated flight time remaining based on most recent lat/lon from ASDI position. I then spent a LOT of time trying to determine heuristics for when one\r\n or more of these were unreliable/erroneous. I threw in a couple variables at the end that measured whether the flight was<br>\r\nprimarily going North, South, East, or West. My final model for runway arrival (combo of SVM and Random Forest) only used 12 variables.&nbsp; Then for gate arrival I used a 3-variable linear regression model to determine runway-to-gate difference.</p>\r\n<p>So it ended up being heavy on data-cleaning/error-detection and light on modeling.</p>\r\n<p>Steve</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22101,
      "author_name": "pprett",
      "author_url": "",
      "post_date": "04/03/2013 19:26:33",
      "content": "<p>@BlackMagic: I used features from all files except the weather forecasts (only current weather) and the flightplans (I did use the last waypoint before the arrival airport).</p>\r\n<p>My overal approach was to use a custom gradient boosting machine that was able to predict multiple outputs; thus, I didn't train individual models for runway and gate arrival or stacked the gate model on top of the runway model. Furthermore, I didn't do\r\n any data cleansing and post-processing of my predicitons. I used the expected arrival benchmark as my initial model, thus my gbm started with the EAB predictions and made&nbsp;</p>\r\n<p>The most important features are derived from the various times (arrival, departure, scheduled, block). Then comes longitude/latitude information of the arrival airport - I think this is used by my model to group arrival airports to certain regions (and to\r\n pool information rather than using individual arrival airports for which data can be quite sparse).</p>\r\n<p>Another pretty strong feature was the minimum flight time given the last known location of the airplane and its cruising speed. I experimented with the approach angle as well but it didn't seem to help much - maybe because the model already uses the ASDI\r\n location as a surogat. I created numerous additional features by computing time differences (e.g. minimum flight time - time to estimated runway arrival).</p>\r\n<p>According to my model, weather information was only marginally useful; the model didn't pick up any signals in ground delay and deicing reports (too sparse data?).</p>\r\n<p>You can find the feature importance plot of the top 50 most important features (relative scores) in the attachment.</p>\r\n<p>I also included some partial dependence plots - I think they are of no use anymore :-/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22111,
      "author_name": "jimthompson",
      "author_url": "",
      "post_date": "04/03/2013 23:06:17",
      "content": "<p>Like all the other posters, congratulations to all the winners.&nbsp; Inspiration for Phase 2.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22147,
      "author_name": "melisgl",
      "author_url": "",
      "post_date": "04/04/2013 09:45:25",
      "content": "<p>The methods of the winners do not seem all that different to me.&nbsp;<span style=\"line-height:1.4em\">Here is mine:</span></p>\r\n<p><span style=\"line-height:1.4em\">Construct a smoothed, gridded speed map indexed by lat/lon/direction from asdi data giving more weight to recent observations. This is much like a gaussian random field. To predict runway arrival, simulate crossing map by\r\n the route in the latest plan. Fall back on ERA, etc if there is no positions or plan data.</span></p>\r\n<p><span style=\"line-height:1.4em\">The speed map seems to pick up the weather factors nicely. Mostly the wind, I guess.</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22148,
      "author_name": "pprett",
      "author_url": "",
      "post_date": "04/04/2013 09:51:29",
      "content": "<p>@Gabor This sounds quite interesting indeed - can you elaborate more?</p>\r\n<p>It also seems to me that most people made predictions in isolation (the prediction of flight A has not influence on the prediciton of flight B) - has anybody incorporated such inter-dependencies?</p>\r\n<p>I added the scheduled arrival rate at certain intervals (30min, 15min) otherwise predictions are totally independent.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22150,
      "author_name": "paweljankiewicz",
      "author_url": "",
      "post_date": "04/04/2013 10:18:42",
      "content": "<p>@Peter: You are right. The flights A and B can be interdependent if they their predicted arrival times coincide in time. To know if they coincide you have to make some initial predictions. So there is a loop hole. In theory you could do an infinite recursion\r\n where the predictions interact with each other: estimations -&gt; predictions -&gt; estimations -&gt; etc.... This is closer to a simulation. Maybe our implementation of this logic was not right but I didn't notice much improvement using interdependent variables.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22153,
      "author_name": "sergeykozub",
      "author_url": "",
      "post_date": "04/04/2013 12:33:25",
      "content": "<p>@Gabor I used an approach very similar to what you described. I used lat/lon/altitude from asdi data giving more weight to recent observations. This worked well for flights near the destination. When falling back to ERA, I applied a correction to the estimate,\r\n depending on runway approach direction, and this adjustment made a good improvement to the score.</p>\r\n<p>@Peter I have some code that makes predictions based on other flights' estimates on the same approach airway. If I have a good estimate for a flight, I'd predict the other flights following it by adding the time to cover the distance between their positions.\r\n This worked well for cases when flights go in circles waiting for the queue to clear.</p>\r\n<p>@Pawel If my model outputs the same time for several flights, it'd add a delay to form up a valid queue with some intervals (typically 1 minute). This gave me only a slight improvement to the score. I haven't used a loop, just doing this once.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22164,
      "author_name": "naokazumizuta",
      "author_url": "",
      "post_date": "04/04/2013 16:22:43",
      "content": "<p>Congratulations to the winners.</p>\r\n<p>I mostly used obvious time features such as difference between expected and scheduled, and last update from actual departure and mean delay of the airports as well as expected arrival (ERA, EGA and estimatedarrival from asdiflightplan). As for asdiposition,\r\n I used only altitude and groundspeed. Based on this thread, I'd better use lat/long. My final model was made by gbm. I used predicted runway arrival as an input to prediction of gate arrival.</p>\r\n<p>Since I started writing code from Feb 1st, I couldn't make decent submission at the time of the first deadline, Feb 8th. My internal validation on PLB at the time of Feb 15th only using information available before 2/8 was about 6.3 ~ 6.4. I didn't expect\r\n I would finish in top 10 since I couldn't exceed the hurdle of 6.0. I was largely motivated by the forum thread</p>\r\n<p>https://www.gequest.com/c/flight/forums/t/3712/do-you-think-is-too-late-to</p>\r\n<p>and Pawel was right. I couldn't finish in the money...</p>\r\n<p>&nbsp;</p>\r\n<p>To kaggle admin:</p>\r\n<p>I'm sorry for creating the thread to try to reveal or discuss the result the other day. I should have read the embargo thread.</p>\r\n<p>&nbsp;</p>\r\n<p>And thank you for hosting great and really fun competition, Ben.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22468,
      "author_name": "jwkvam",
      "author_url": "",
      "post_date": "04/09/2013 05:27:19",
      "content": "<p>Compared with the approaches of the other competitors I'm afraid my approach was quite plain jane. I simply had two independent gradient boosted regressors for the gate and runway predictions and fed them the same features. I spent most of my time trying\r\n to think of creative features, in the end I had about 1900 total features of which only 1100 were used by the estimator. I used virtually all the available data with the exception of the ATSCC data. I didn't spend much time doing feature selection/optimization.\r\n Instead I tried to feed the algorithms as much training data as possible. In the end I wanted to use 310,000 training flights but I could only handle 260,000 in memory at one time so that's what I went with. I didn't spend much time trying to deal with outliers,\r\n I just used huber loss with a very high alpha value.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22483,
      "author_name": "xavierconort",
      "author_url": "",
      "post_date": "04/09/2013 10:18:11",
      "content": "<p>Like Jacques, we decided to discard our features generated from ATSCC data as we didn't find them useful at the early stage of the competition (too few events).</p>\r\n<p>Our solution is a blend of GBMs and RFs but a simpler solution would have provided us with similar accuracy: 7.987. This solution consists of a GBM to model the ERA error and a RF to model the taxi time (difference between runway and gate arrival times).\r\n Our nb of training examples is similar to Jacques' nb.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22527,
      "author_name": "gtakacs",
      "author_url": "",
      "post_date": "04/09/2013 21:11:46",
      "content": "<p>I used a relatively small number of features (56) and tried to squeeze out as much as possible from them. I generated a random cutoff time for each flight in order to have more training examples. I trained my final model on 850,000 flights (split into 3\r\n parts).</p>\r\n<p>The most informative features in my model were direct estimates of the arrival times (with some cleaning applied). The second best ones were categorical features associated with the flights (e.g. destination, airline, destination x airline, destination x\r\n origin, etc.). I ended up with implementing my own sparse ridge regression solver, because scikit-learn's did not converge on this data.</p>\r\n<p>I also used weather and positional features. The most useful weather feature was the destination airport's most recent &quot;present conditions&quot; value.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 35948,
      "author_name": "",
      "author_url": "",
      "post_date": "12/07/2013 15:47:33",
      "content": "<p>I want to analyze this data for education process. Can anyone help on how to start the process of cleaning the data and suggest some techniques.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "22058": "",
    "22060": "",
    "22061": "",
    "22062": "",
    "22064": "",
    "22067": "",
    "22071": "",
    "22072": "",
    "22079": "",
    "22083": "",
    "22088": "",
    "22090": "",
    "22099": "",
    "22101": "",
    "22111": "",
    "22147": "",
    "22148": "",
    "22150": "",
    "22153": "",
    "22164": "",
    "22468": "",
    "22483": "",
    "22527": "",
    "35948": ""
  },
  "source": "meta"
}