{
  "id": 3234,
  "title": "Negative taxi times",
  "url": "/competitions/flight/discussion/3234",
  "author_name": "",
  "post_date": "2012-12-02T15:36:25.827Z",
  "votes": 3,
  "comment_count": 2,
  "views": 3444,
  "content": "<p>I've found an interresting issue here. There are 941 records (3.5% of the data)&nbsp; in the sample dataset where the actual_gate_arrival is earlier than the&nbsp;actual_runway_arrival. Which means that the plane gets to the gate first and after than lands on the\r\n runway. That would be quite a spectacular maneuver from a hundred-tons jetliner however I rather see it as a data quality error. By looking at a specific example flight_history_id=280878607 (JetBlue 493 from BOS to DEN), you will find the following rows in\r\n flighthistoryevents.csv:</p>\r\n<p>280878607,2012-11-20 10:14:19.107-08,Time Adjustment,ARA- New=11/20/12 11:06<br>\r\n280878607,2012-11-20 10:13:51.45-08,STATUS-Landed,&quot;AGA- New=11/20/12 11:01, STATUS- Old=A New=L&quot;</p>\r\n<p>which states that landing happened 5 mins after the actual gate arrival. By looking up this flight in the asdiposition.csv file:</p>\r\n<p>2012-11-20 09:54:35-08,JBU493,5900,140,39.9199981689453,-104.680000305176,280878607<br>\r\n2012-11-20 09:54:55-08,JBU493,5700,139,39.9000015258789,-104.680000305176,280878607<br>\r\n2012-11-20 09:55:57-08,JBU493,5400,113,39.8699989318848,-104.680000305176,280878607<br>\r\n2012-11-20 09:56:59-08,JBU493,5400,50,39.8699989318848,-104.680000305176,280878607</p>\r\n<p>that shows that by 10:56 local the plane was already on the ground (Denver airport has an elevation of 5433 ft). So the real landing happened 10 mins earlier than the &quot;actual&quot; landing time given in the flight history.</p>\r\n<p>The big problem here is that these variables (actual_gate_arrival, actual_runway_arrival) are our target variables and assuming similar data quality errors in the evaluation dataset would mean that we should create negative taxi time predictions. Technically\r\n it is attainable&nbsp; but from business point of view it is nonsensical. With other words: our models will not&nbsp; focus 100 percent on predicting gate/runway arrival times but will predict partially the data quality errors. 3.5% of data has obviously erroneous AGA\r\n or ARA values (and maybe much more if someone would do a more in-depth analysis), so it could introduce a significant bias into the model!</p>\r\n<p>IMHO the comp organizer should consider:</p>\r\n<p>a. entirely omitting flights with erroneous arrival values from both the training and evalutaion dataset<br>\r\nor <br>\r\nb. creating a weight variable which measures the reliability of the target values. This weight could be used when calculating the submission RMSE. By publishing it for the training data we will able to tune our models so that they will focus on reliable records.</p>",
  "messages": [
    {
      "id": "17343",
      "postDate": "12/02/2012 15:36:25",
      "content": "<p>I've found an interresting issue here. There are 941 records (3.5% of the data)&nbsp; in the sample dataset where the actual_gate_arrival is earlier than the&nbsp;actual_runway_arrival. Which means that the plane gets to the gate first and after than lands on the\r\n runway. That would be quite a spectacular maneuver from a hundred-tons jetliner however I rather see it as a data quality error. By looking at a specific example flight_history_id=280878607 (JetBlue 493 from BOS to DEN), you will find the following rows in\r\n flighthistoryevents.csv:</p>\r\n<p>280878607,2012-11-20 10:14:19.107-08,Time Adjustment,ARA- New=11/20/12 11:06<br>\r\n280878607,2012-11-20 10:13:51.45-08,STATUS-Landed,&quot;AGA- New=11/20/12 11:01, STATUS- Old=A New=L&quot;</p>\r\n<p>which states that landing happened 5 mins after the actual gate arrival. By looking up this flight in the asdiposition.csv file:</p>\r\n<p>2012-11-20 09:54:35-08,JBU493,5900,140,39.9199981689453,-104.680000305176,280878607<br>\r\n2012-11-20 09:54:55-08,JBU493,5700,139,39.9000015258789,-104.680000305176,280878607<br>\r\n2012-11-20 09:55:57-08,JBU493,5400,113,39.8699989318848,-104.680000305176,280878607<br>\r\n2012-11-20 09:56:59-08,JBU493,5400,50,39.8699989318848,-104.680000305176,280878607</p>\r\n<p>that shows that by 10:56 local the plane was already on the ground (Denver airport has an elevation of 5433 ft). So the real landing happened 10 mins earlier than the &quot;actual&quot; landing time given in the flight history.</p>\r\n<p>The big problem here is that these variables (actual_gate_arrival, actual_runway_arrival) are our target variables and assuming similar data quality errors in the evaluation dataset would mean that we should create negative taxi time predictions. Technically\r\n it is attainable&nbsp; but from business point of view it is nonsensical. With other words: our models will not&nbsp; focus 100 percent on predicting gate/runway arrival times but will predict partially the data quality errors. 3.5% of data has obviously erroneous AGA\r\n or ARA values (and maybe much more if someone would do a more in-depth analysis), so it could introduce a significant bias into the model!</p>\r\n<p>IMHO the comp organizer should consider:</p>\r\n<p>a. entirely omitting flights with erroneous arrival values from both the training and evalutaion dataset<br>\r\nor <br>\r\nb. creating a weight variable which measures the reliability of the target values. This weight could be used when calculating the submission RMSE. By publishing it for the training data we will able to tune our models so that they will focus on reliable records.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17363",
      "postDate": "12/03/2012 14:47:14",
      "content": "<p>This is a pretty big problem since we are trying to predict the arrival times.&nbsp; What's troubling is that until we can identify the cause of this error, we don't know what other errors there may be which will only serve to compound the problem.&nbsp; Maybe this\r\n is a simple matter that not all airport systems log and track the data correctly, or maybe the ILS system was down and this data was logged by another system, etc, etc.&nbsp;&nbsp; I'm only guessing here, so does anyone have knowledge of how these times are collected?&nbsp;\r\n We'd need someone with knowlege of the acutal systems to give some insight.&nbsp; This may help develop a weight metric like you mentioned.</p>\r\n<p>Obviously there is a work around, which is to use only the flight data to determine arriaval times, but that also means an extra step of correlating runway lat/long/elevation to dermine the exact arrival.&nbsp; More than that though, what if the plane's sitting\r\n on the taxiway waiting to dock?&nbsp; How do we know the actual gate time?</p>\r\n<p>Ah the joys of real-world data ;)</p>\r\n<p>&nbsp;</p>\r\n<p>EDIT :</p>\r\n<p>Another thought on this issue,&nbsp; how do we know if the final test data will be free from these errors?&nbsp; Has someone manually verified them?&nbsp; If we are using the erroneous data to test the final submissions, the workaround I mentioned above might actually\r\n cause the submission to score lower than it should even though it's actually more acurate in the real-world than one which does not.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "17364",
      "postDate": "12/03/2012 15:28:38",
      "content": "<p>The data set is extremely real life. Do they type in everything by hand? Inconsistencies and outright mistakes abound.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 17363,
      "author_name": "mcstar",
      "author_url": "",
      "post_date": "12/03/2012 14:47:14",
      "content": "<p>This is a pretty big problem since we are trying to predict the arrival times.&nbsp; What's troubling is that until we can identify the cause of this error, we don't know what other errors there may be which will only serve to compound the problem.&nbsp; Maybe this\r\n is a simple matter that not all airport systems log and track the data correctly, or maybe the ILS system was down and this data was logged by another system, etc, etc.&nbsp;&nbsp; I'm only guessing here, so does anyone have knowledge of how these times are collected?&nbsp;\r\n We'd need someone with knowlege of the acutal systems to give some insight.&nbsp; This may help develop a weight metric like you mentioned.</p>\r\n<p>Obviously there is a work around, which is to use only the flight data to determine arriaval times, but that also means an extra step of correlating runway lat/long/elevation to dermine the exact arrival.&nbsp; More than that though, what if the plane's sitting\r\n on the taxiway waiting to dock?&nbsp; How do we know the actual gate time?</p>\r\n<p>Ah the joys of real-world data ;)</p>\r\n<p>&nbsp;</p>\r\n<p>EDIT :</p>\r\n<p>Another thought on this issue,&nbsp; how do we know if the final test data will be free from these errors?&nbsp; Has someone manually verified them?&nbsp; If we are using the erroneous data to test the final submissions, the workaround I mentioned above might actually\r\n cause the submission to score lower than it should even though it's actually more acurate in the real-world than one which does not.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 17364,
      "author_name": "melisgl",
      "author_url": "",
      "post_date": "12/03/2012 15:28:38",
      "content": "<p>The data set is extremely real life. Do they type in everything by hand? Inconsistencies and outright mistakes abound.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "17343": "",
    "17363": "",
    "17364": ""
  },
  "source": "meta"
}