{
  "id": 3555,
  "title": "Thoughts on the Solution",
  "url": "/competitions/flight/discussion/3555",
  "author_name": "",
  "post_date": "2013-01-08T15:09:42.517Z",
  "votes": null,
  "comment_count": 4,
  "views": 3920,
  "content": "<p>If one was to start from an error of about 9 minutes (achievable by a straightforward use of the estimates with corrections for bad data), and with a minimum achievable error of about 5 by an estimate of noise, I think you will find that of the four minutes\r\n that can be lost, 2 is from ETA and 2 is from the error in taxi times. The error reductions in both cases are almost entirely due to the long tails of the respective Poisson distributions.</p>\r\n<p>The problem comes down to modeling the exceptions (mostly due to congestion), and from good unbiased estimates for each of the particular airports/terminals and airlines. Once you've achieved an error about 5 minutes, the trick then seems to be to ensure\r\n that the predictions for the final test data are temporally aligned (i.e. Monday's parameters for a Monday of test, Saturday for Saturday, morning for morning), and stay away from using data from odd events i.e. O'Hare during the Thanksgiving rush. If you\r\n were to lift the data temporally, and by airline perhaps for the larger airports, then again I think you will see that the problem comes down to brute force modeling of individual terminals.</p>\r\n<p>So, of the various participants, we ought to have at least seven teams that can reduce this problem to chasing errors at the flight level by the time all is said and done. At that point it is just a matter of luck which team will win within the noise level.\r\n I expect the team with the better models of the larger airports will tend to come out on top.</p>\r\n<p>Agree/Disagree? This seems like a data driven modeling problem to me (i.e. not so much of a data mining problem).</p>\r\n<p>&nbsp;</p>\r\n<p>Oh, I should add: The one thing not to do is the typical massive cross-validation. If you do this using the whole data set, you will pick up the aberrations of the holiday congestion (which does not apply to February). Alternatively, you may not be selective\r\n enough in the treatment of the tails since there is not so much data once you are done lifting it by airline/airport and temporally (and this will produce high variance in the estimates). So there may be some participants near the top of the leaderboard that\r\n will have a bad go of it on the final data set.</p>",
  "messages": [
    {
      "id": "19093",
      "postDate": "01/08/2013 15:09:42",
      "content": "<p>If one was to start from an error of about 9 minutes (achievable by a straightforward use of the estimates with corrections for bad data), and with a minimum achievable error of about 5 by an estimate of noise, I think you will find that of the four minutes\r\n that can be lost, 2 is from ETA and 2 is from the error in taxi times. The error reductions in both cases are almost entirely due to the long tails of the respective Poisson distributions.</p>\r\n<p>The problem comes down to modeling the exceptions (mostly due to congestion), and from good unbiased estimates for each of the particular airports/terminals and airlines. Once you've achieved an error about 5 minutes, the trick then seems to be to ensure\r\n that the predictions for the final test data are temporally aligned (i.e. Monday's parameters for a Monday of test, Saturday for Saturday, morning for morning), and stay away from using data from odd events i.e. O'Hare during the Thanksgiving rush. If you\r\n were to lift the data temporally, and by airline perhaps for the larger airports, then again I think you will see that the problem comes down to brute force modeling of individual terminals.</p>\r\n<p>So, of the various participants, we ought to have at least seven teams that can reduce this problem to chasing errors at the flight level by the time all is said and done. At that point it is just a matter of luck which team will win within the noise level.\r\n I expect the team with the better models of the larger airports will tend to come out on top.</p>\r\n<p>Agree/Disagree? This seems like a data driven modeling problem to me (i.e. not so much of a data mining problem).</p>\r\n<p>&nbsp;</p>\r\n<p>Oh, I should add: The one thing not to do is the typical massive cross-validation. If you do this using the whole data set, you will pick up the aberrations of the holiday congestion (which does not apply to February). Alternatively, you may not be selective\r\n enough in the treatment of the tails since there is not so much data once you are done lifting it by airline/airport and temporally (and this will produce high variance in the estimates). So there may be some participants near the top of the leaderboard that\r\n will have a bad go of it on the final data set.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19103",
      "postDate": "01/08/2013 19:03:48",
      "content": "<p>pdeignan,<br>\r\nGood thoughts.<br>\r\nHow are you doing cross-validation or how do you recommend it?</p>\r\n<p>I see that just taking exp gate arrival mins and exp runway arrival mins gives a score of 7.2 for runway and 8.5 for gate but actually the score is 9.23 on leaderboard.<br>\r\nI think a good CV mechanism will be compulsory for this</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19108",
      "postDate": "01/08/2013 19:33:07",
      "content": "<p>The problem is so low-dim that I would not recommend cross validation except in certain exploratory and well controlled situations.</p>\r\n<p>For example, take a look at flights into Nantucket (ATL is also a beautiful example). There is a typical Poisson distribution for the various airlines (and note that the airlines here have different mean taxi times). You will see that once you account for\r\n the tails (to what degree they can be predicted by an in depth examination of the particular events) the means are easily determined. Of course, you might look also at the time of day and day of week as factors.</p>\r\n<p>Either way, besides for the tails, this is a sequential one or two dim estimate at best. So do this for each airport (80% or so are just executives flights for the holidays into single strip airports with virtually no other traffic, i.e. the times for these\r\n airports can be taken by immediate inspection.)</p>\r\n<p>While the leaderboard results might look great for massive cross validation (as they always will for small datasets), you will be making a fatal mistake by approaching the problem this way.</p>\r\n<p>Decompose it first by airport. ETA and taxi times are virtually seperate problems. You may look at congestion as the greatest cause of ETAs being other than posted (the ETA is actually a closed loop tracking problem--pilots try to make their ETA). Taxi time\r\n tails are also likely due to congestion at the gate (but you will need to look at these cases individually being very clear not to create false positives in your model.</p>\r\n<p>&nbsp;</p>\r\n<p>Bottom line: this is a modeling problem.</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19140",
      "postDate": "01/10/2013 02:30:27",
      "content": "<p>Hi Pdeigan,</p>\r\n<p>You seem upset that you can't just throw this into a random forrest and get an answer. I like that this is more modeling that machine learning and I hope there are more contests like this. Something I've been thinking about, but haven't had time to implement\r\n is trying to estimate the time remaining in a flight. We have the departure time, we know where the flight is headed and we know at the cutoff time the GPS position of the aircraft. Seems to me like we should be able to figure estimate the time remaining,\r\n understanding that it's not linear and that there are flight phases like ascent, cruise, and descent.</p>\r\n<p>Anyways, thanks for sharing your ideas. I had been examining the airports individually but hadn't thought to go as far as you. I split it out by hour, but completely ignored that day of week matters. I also never considered that airline will matter.</p>\r\n<p>Last thing, you kept confusing me when you used the word &quot;tails.&quot; To some people &quot;tails&quot; are aircraft. But I think you are refering to probability distributions :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19144",
      "postDate": "01/10/2013 03:25:23",
      "content": "<p>Au contraire Ricardo,</p>\r\n<p>I'm no fan of brainless methods. They have their places, but anyone with a computer can do them.</p>\r\n<p>It is clear that the problem can be modeled as (and what would one expect???) ..... a simple parameterized queuing model. This is the conclusion from the observation of Poisson behavior of the distributions. The respective independent variables can be calculated\r\n (the rules for these queues are well known to airline travelers) and parameters fit.</p>\r\n<p>Lovely, but big deal. Where is the suspense and mystery in this?</p>\r\n<p>How many people in the transportation industry do you think have very good models of congested transportation networks? (My guess: all) Is the contribution of this effort breaking new ground? (my guess is no--not in this limited time and unless of course\r\n we made this a real competition).</p>\r\n<p>What do I mean by that? Well, first, by making it a competition amongst equals and of such complexity that the clever insight and the deep thought might yield the final winner in a tough fight. Perhaps I was the only one bored to death of the method of solution\r\n of the Netflix prize. Notice that for a million dollars no one is running to AT&amp;T to put their algorithms in industry standard software other than the particular case of Netflix.</p>\r\n<p>With so many headed down the overfitting path with CV (and there are not that many participants for the money offered), my hopes were that we could step it up a notch. However, the problem still decomposes to a moderate number of simple problems . Where\r\n is the fun in that? It would make more sense to begin with some knowledge of the present methods and build from there but I suppose there is no one centralized deconfliction authority. (Is that true???).</p>\r\n<p>Yes, tails -- they are from a nonlinear queue (saturation effects). But this too can be parameterized easily from the data.</p>\r\n<p>BTW, chose carefully your independent parameters. Time of day may be highly correlated, but is that really it? It should be the number in the queues. Of course, you have to choose the queues, but this is easy from an examination of the data such as I have\r\n plotted.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 19103,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "01/08/2013 19:03:48",
      "content": "<p>pdeignan,<br>\r\nGood thoughts.<br>\r\nHow are you doing cross-validation or how do you recommend it?</p>\r\n<p>I see that just taking exp gate arrival mins and exp runway arrival mins gives a score of 7.2 for runway and 8.5 for gate but actually the score is 9.23 on leaderboard.<br>\r\nI think a good CV mechanism will be compulsory for this</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 19108,
      "author_name": "pdeignan",
      "author_url": "",
      "post_date": "01/08/2013 19:33:07",
      "content": "<p>The problem is so low-dim that I would not recommend cross validation except in certain exploratory and well controlled situations.</p>\r\n<p>For example, take a look at flights into Nantucket (ATL is also a beautiful example). There is a typical Poisson distribution for the various airlines (and note that the airlines here have different mean taxi times). You will see that once you account for\r\n the tails (to what degree they can be predicted by an in depth examination of the particular events) the means are easily determined. Of course, you might look also at the time of day and day of week as factors.</p>\r\n<p>Either way, besides for the tails, this is a sequential one or two dim estimate at best. So do this for each airport (80% or so are just executives flights for the holidays into single strip airports with virtually no other traffic, i.e. the times for these\r\n airports can be taken by immediate inspection.)</p>\r\n<p>While the leaderboard results might look great for massive cross validation (as they always will for small datasets), you will be making a fatal mistake by approaching the problem this way.</p>\r\n<p>Decompose it first by airport. ETA and taxi times are virtually seperate problems. You may look at congestion as the greatest cause of ETAs being other than posted (the ETA is actually a closed loop tracking problem--pilots try to make their ETA). Taxi time\r\n tails are also likely due to congestion at the gate (but you will need to look at these cases individually being very clear not to create false positives in your model.</p>\r\n<p>&nbsp;</p>\r\n<p>Bottom line: this is a modeling problem.</p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 19140,
      "author_name": "rickyars",
      "author_url": "",
      "post_date": "01/10/2013 02:30:27",
      "content": "<p>Hi Pdeigan,</p>\r\n<p>You seem upset that you can't just throw this into a random forrest and get an answer. I like that this is more modeling that machine learning and I hope there are more contests like this. Something I've been thinking about, but haven't had time to implement\r\n is trying to estimate the time remaining in a flight. We have the departure time, we know where the flight is headed and we know at the cutoff time the GPS position of the aircraft. Seems to me like we should be able to figure estimate the time remaining,\r\n understanding that it's not linear and that there are flight phases like ascent, cruise, and descent.</p>\r\n<p>Anyways, thanks for sharing your ideas. I had been examining the airports individually but hadn't thought to go as far as you. I split it out by hour, but completely ignored that day of week matters. I also never considered that airline will matter.</p>\r\n<p>Last thing, you kept confusing me when you used the word &quot;tails.&quot; To some people &quot;tails&quot; are aircraft. But I think you are refering to probability distributions :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 19144,
      "author_name": "pdeignan",
      "author_url": "",
      "post_date": "01/10/2013 03:25:23",
      "content": "<p>Au contraire Ricardo,</p>\r\n<p>I'm no fan of brainless methods. They have their places, but anyone with a computer can do them.</p>\r\n<p>It is clear that the problem can be modeled as (and what would one expect???) ..... a simple parameterized queuing model. This is the conclusion from the observation of Poisson behavior of the distributions. The respective independent variables can be calculated\r\n (the rules for these queues are well known to airline travelers) and parameters fit.</p>\r\n<p>Lovely, but big deal. Where is the suspense and mystery in this?</p>\r\n<p>How many people in the transportation industry do you think have very good models of congested transportation networks? (My guess: all) Is the contribution of this effort breaking new ground? (my guess is no--not in this limited time and unless of course\r\n we made this a real competition).</p>\r\n<p>What do I mean by that? Well, first, by making it a competition amongst equals and of such complexity that the clever insight and the deep thought might yield the final winner in a tough fight. Perhaps I was the only one bored to death of the method of solution\r\n of the Netflix prize. Notice that for a million dollars no one is running to AT&amp;T to put their algorithms in industry standard software other than the particular case of Netflix.</p>\r\n<p>With so many headed down the overfitting path with CV (and there are not that many participants for the money offered), my hopes were that we could step it up a notch. However, the problem still decomposes to a moderate number of simple problems . Where\r\n is the fun in that? It would make more sense to begin with some knowledge of the present methods and build from there but I suppose there is no one centralized deconfliction authority. (Is that true???).</p>\r\n<p>Yes, tails -- they are from a nonlinear queue (saturation effects). But this too can be parameterized easily from the data.</p>\r\n<p>BTW, chose carefully your independent parameters. Time of day may be highly correlated, but is that really it? It should be the number in the queues. Of course, you have to choose the queues, but this is easy from an examination of the data such as I have\r\n plotted.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "19093": "",
    "19103": "",
    "19108": "",
    "19140": "",
    "19144": ""
  },
  "source": "meta"
}