{
  "id": 3477,
  "title": "Extra entries when trying to generate \"test_flights_combined.csv\"",
  "url": "/competitions/flight/discussion/3477",
  "author_name": "",
  "post_date": "2012-12-27T01:02:26.013Z",
  "votes": null,
  "comment_count": 8,
  "views": 5274,
  "content": "<pre>Hi,<br><br>I've been trying to generate my own version of the &quot;test_flights_combined&quot; file, with some additional pre-calculated columns, but for some reason I keep on getting more entries than the original file published in the Data page.<br>The &quot;test_flights_combined&quot; and the benchmark files contain 26459 entries, but I end up with 27786.<br>I tried to use the same algorithm as in the flighthistory.py code I found in GitHub (the &quot;flight_history_row_in_test_set&quot; method, to be precise) to determine whether the flight is in the air or not. Although I ported the Python code to the language I'm using, it should be fairly similar.<br><br>Below, you can see a few examples of the extra flights I'm getting, and by manually checking the conditions (cut off time, continental US flights, missing values) they look valid to me, but weren't included in the test set.<br><br>281396084,WN,SWA,3083,BDL,KBDL,DEN,KDEN,2012-11-26 13:30:00&#43;00:00,2012-11-26 18:05:00&#43;00:00,2012-11-26 13:30:00&#43;00:00,2012-11-26 13:38:00&#43;00:00,2012-11-26 18:05:00&#43;00:00,HIDDEN,2012-11-26 13:40:00&#43;00:00,2012-11-26 13:45:00&#43;00:00,2012-11-26 17:54:00&#43;00:00,HIDDEN,I,254,275,-5,-7,73W,,B737<br><br>281908281,WN,SWA,3170,SNA,KSNA,DEN,KDEN,2012-12-01 21:55:00&#43;00:00,2012-12-02 00:05:00&#43;00:00,2012-12-01 21:55:00&#43;00:00,2012-12-01 22:11:00&#43;00:00,2012-12-02 00:05:00&#43;00:00,HIDDEN,2012-12-01 22:07:00&#43;00:00,2012-12-01 22:19:00&#43;00:00,2012-12-01 23:53:00&#43;00:00,HIDDEN,I,106,130,-8,-7,73G,,B737<br><br>282239833,OO,SKW,6219,MEM,KMEM,DEN,KDEN,2012-12-05 20:22:00&#43;00:00,2012-12-05 22:59:00&#43;00:00,2012-12-05 20:22:00&#43;00:00,2012-12-05 20:12:00&#43;00:00,2012-12-05 22:59:00&#43;00:00,HIDDEN,2012-12-05 20:32:00&#43;00:00,2012-12-05 20:29:00&#43;00:00,2012-12-05 22:35:00&#43;00:00,HIDDEN,I,123,157,-6,-7,CR7,,CRJ7</pre>",
  "messages": [
    {
      "id": "18651",
      "postDate": "12/27/2012 01:02:26",
      "content": "<pre>Hi,<br><br>I've been trying to generate my own version of the &quot;test_flights_combined&quot; file, with some additional pre-calculated columns, but for some reason I keep on getting more entries than the original file published in the Data page.<br>The &quot;test_flights_combined&quot; and the benchmark files contain 26459 entries, but I end up with 27786.<br>I tried to use the same algorithm as in the flighthistory.py code I found in GitHub (the &quot;flight_history_row_in_test_set&quot; method, to be precise) to determine whether the flight is in the air or not. Although I ported the Python code to the language I'm using, it should be fairly similar.<br><br>Below, you can see a few examples of the extra flights I'm getting, and by manually checking the conditions (cut off time, continental US flights, missing values) they look valid to me, but weren't included in the test set.<br><br>281396084,WN,SWA,3083,BDL,KBDL,DEN,KDEN,2012-11-26 13:30:00&#43;00:00,2012-11-26 18:05:00&#43;00:00,2012-11-26 13:30:00&#43;00:00,2012-11-26 13:38:00&#43;00:00,2012-11-26 18:05:00&#43;00:00,HIDDEN,2012-11-26 13:40:00&#43;00:00,2012-11-26 13:45:00&#43;00:00,2012-11-26 17:54:00&#43;00:00,HIDDEN,I,254,275,-5,-7,73W,,B737<br><br>281908281,WN,SWA,3170,SNA,KSNA,DEN,KDEN,2012-12-01 21:55:00&#43;00:00,2012-12-02 00:05:00&#43;00:00,2012-12-01 21:55:00&#43;00:00,2012-12-01 22:11:00&#43;00:00,2012-12-02 00:05:00&#43;00:00,HIDDEN,2012-12-01 22:07:00&#43;00:00,2012-12-01 22:19:00&#43;00:00,2012-12-01 23:53:00&#43;00:00,HIDDEN,I,106,130,-8,-7,73G,,B737<br><br>282239833,OO,SKW,6219,MEM,KMEM,DEN,KDEN,2012-12-05 20:22:00&#43;00:00,2012-12-05 22:59:00&#43;00:00,2012-12-05 20:22:00&#43;00:00,2012-12-05 20:12:00&#43;00:00,2012-12-05 22:59:00&#43;00:00,HIDDEN,2012-12-05 20:32:00&#43;00:00,2012-12-05 20:29:00&#43;00:00,2012-12-05 22:35:00&#43;00:00,HIDDEN,I,123,157,-6,-7,CR7,,CRJ7</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18654",
      "postDate": "12/27/2012 03:54:27",
      "content": "<p>This happens with my Python code too. Can't figure out what the issue is either.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18663",
      "postDate": "12/27/2012 11:49:17",
      "content": "<p>I didn't try to generate own versions of the test_flights file yet, but I think a possible reason may be that not all data is visible to us. For example it was mentioned somere here in the board that only flights with a positive (or zero) actual gate arrival\r\n and actual runway arrival difference are selected. I think this is because predicting negative taxi times contradicts everything related to reality ;) Since this informations are missing for us (wer have to predict them) it feels quite natural that self-written\r\n filters may pass more flights than the official ones.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18689",
      "postDate": "12/27/2012 21:56:29",
      "content": "<p>Edit: Nevermind on the getting test flights from github, im just going to use the benchmark IDs to know what i need to prediction for.</p>\r\n<p>For the final set can we assume that we will get the IDs of the flights which we should generate predictions for?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18825",
      "postDate": "12/31/2012 04:02:05",
      "content": "<p>[quote=David Leen;18654]</p>\r\n<p>This happens with my Python code too. Can't figure out what the issue is either.</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>Did you reach the same number as Gianluca? My matlab code ended up with 27627 rather than 27786...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18938",
      "postDate": "01/03/2013 17:49:19",
      "content": "<p>After adding in an additional constraint - no diverted or re-directed flights - I get 27,872 to predict. I should note that I also assumed no cancelled flights, even though they weren't specifically addressed in this post:&nbsp;https://www.gequest.com/c/flight/forums/t/3520/redirected-diverted-flights</p>\r\n<p>Are we going to get clarificaiton on how to determine the test flights from the data? Or can we assume that we'll always be given the &quot;test_flights_combined&quot; file from which we can pull the ids? I think we have to be given the ids since we can't check the\r\n gate arrival &lt; runway arrival constraint.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19112",
      "postDate": "01/08/2013 19:43:26",
      "content": "<p>[quote=ricardo sachez;18938]</p>\r\n<p>After adding in an additional constraint - no diverted or re-directed flights - I get 27,872 to predict. I should note that I also assumed no cancelled flights, even though they weren't specifically addressed in this post:&nbsp;https://www.gequest.com/c/flight/forums/t/3520/redirected-diverted-flights</p>\r\n<p>Are we going to get clarificaiton on how to determine the test flights from the data? Or can we assume that we'll always be given the &quot;test_flights_combined&quot; file from which we can pull the ids? I think we have to be given the ids since we can't check the\r\n gate arrival &lt; runway arrival constraint.</p>\r\n<p>[/quote]We will always provide the list of ids for the flights you are to predict. As you pointed out, you can't determine this list yourself since you can't check all the constraints in the public leaderboard data (e.g. gate_arrival &lt; runway_arrival).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19963",
      "postDate": "02/03/2013 07:15:12",
      "content": "<p>To confirm that I understand:</p>\r\n<p>First, we submit a file that can produce our final answers.</p>\r\n<p>Second, the final data is posted. We push the button, get the answers, and submit the answers.</p>\r\n<p>Third, Kaggle ranks the answers and checks the submitted file for consistency.</p>\r\n<p>My test solution set also has too many entries (28,313, about 2,000 too many). If I understand correctly, I can clean the extra answers from the solution file that I submit. However, if Kaggle uses my program, it will generate some extra answers which will\r\n be ignored(?) Thanks,</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19965",
      "postDate": "02/03/2013 08:53:45",
      "content": "<p>[quote=Michael Bickel;19963]</p>\r\n<p>To confirm that I understand:</p>\r\n<p>First, we submit a file that can produce our final answers.</p>\r\n<p>Second, the final data is posted. We push the button, get the answers, and submit the answers.</p>\r\n<p>Third, Kaggle ranks the answers and checks the submitted file for consistency.</p>\r\n<p>[/quote]This is correct</p>\r\n<p>[quote=Michael Bickel;19963]My test solution set also has too many entries (28,313, about 2,000 too many). If I understand correctly, I can clean the extra answers from the solution file that I submit. However, if Kaggle uses my program, it will generate\r\n some extra answers which will be ignored(?) Thanks,[/quote]Removing any extra entries should be automated &amp; part of your model code.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 18654,
      "author_name": "davidleen",
      "author_url": "",
      "post_date": "12/27/2012 03:54:27",
      "content": "<p>This happens with my Python code too. Can't figure out what the issue is either.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18663,
      "author_name": "lolliedieb",
      "author_url": "",
      "post_date": "12/27/2012 11:49:17",
      "content": "<p>I didn't try to generate own versions of the test_flights file yet, but I think a possible reason may be that not all data is visible to us. For example it was mentioned somere here in the board that only flights with a positive (or zero) actual gate arrival\r\n and actual runway arrival difference are selected. I think this is because predicting negative taxi times contradicts everything related to reality ;) Since this informations are missing for us (wer have to predict them) it feels quite natural that self-written\r\n filters may pass more flights than the official ones.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18689,
      "author_name": "miketery",
      "author_url": "",
      "post_date": "12/27/2012 21:56:29",
      "content": "<p>Edit: Nevermind on the getting test flights from github, im just going to use the benchmark IDs to know what i need to prediction for.</p>\r\n<p>For the final set can we assume that we will get the IDs of the flights which we should generate predictions for?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18825,
      "author_name": "jay670790",
      "author_url": "",
      "post_date": "12/31/2012 04:02:05",
      "content": "<p>[quote=David Leen;18654]</p>\r\n<p>This happens with my Python code too. Can't figure out what the issue is either.</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>Did you reach the same number as Gianluca? My matlab code ended up with 27627 rather than 27786...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18938,
      "author_name": "rickyars",
      "author_url": "",
      "post_date": "01/03/2013 17:49:19",
      "content": "<p>After adding in an additional constraint - no diverted or re-directed flights - I get 27,872 to predict. I should note that I also assumed no cancelled flights, even though they weren't specifically addressed in this post:&nbsp;https://www.gequest.com/c/flight/forums/t/3520/redirected-diverted-flights</p>\r\n<p>Are we going to get clarificaiton on how to determine the test flights from the data? Or can we assume that we'll always be given the &quot;test_flights_combined&quot; file from which we can pull the ids? I think we have to be given the ids since we can't check the\r\n gate arrival &lt; runway arrival constraint.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 19112,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "01/08/2013 19:43:26",
      "content": "<p>[quote=ricardo sachez;18938]</p>\r\n<p>After adding in an additional constraint - no diverted or re-directed flights - I get 27,872 to predict. I should note that I also assumed no cancelled flights, even though they weren't specifically addressed in this post:&nbsp;https://www.gequest.com/c/flight/forums/t/3520/redirected-diverted-flights</p>\r\n<p>Are we going to get clarificaiton on how to determine the test flights from the data? Or can we assume that we'll always be given the &quot;test_flights_combined&quot; file from which we can pull the ids? I think we have to be given the ids since we can't check the\r\n gate arrival &lt; runway arrival constraint.</p>\r\n<p>[/quote]We will always provide the list of ids for the flights you are to predict. As you pointed out, you can't determine this list yourself since you can't check all the constraints in the public leaderboard data (e.g. gate_arrival &lt; runway_arrival).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 19963,
      "author_name": "michaelbickel",
      "author_url": "",
      "post_date": "02/03/2013 07:15:12",
      "content": "<p>To confirm that I understand:</p>\r\n<p>First, we submit a file that can produce our final answers.</p>\r\n<p>Second, the final data is posted. We push the button, get the answers, and submit the answers.</p>\r\n<p>Third, Kaggle ranks the answers and checks the submitted file for consistency.</p>\r\n<p>My test solution set also has too many entries (28,313, about 2,000 too many). If I understand correctly, I can clean the extra answers from the solution file that I submit. However, if Kaggle uses my program, it will generate some extra answers which will\r\n be ignored(?) Thanks,</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 19965,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "02/03/2013 08:53:45",
      "content": "<p>[quote=Michael Bickel;19963]</p>\r\n<p>To confirm that I understand:</p>\r\n<p>First, we submit a file that can produce our final answers.</p>\r\n<p>Second, the final data is posted. We push the button, get the answers, and submit the answers.</p>\r\n<p>Third, Kaggle ranks the answers and checks the submitted file for consistency.</p>\r\n<p>[/quote]This is correct</p>\r\n<p>[quote=Michael Bickel;19963]My test solution set also has too many entries (28,313, about 2,000 too many). If I understand correctly, I can clean the extra answers from the solution file that I submit. However, if Kaggle uses my program, it will generate\r\n some extra answers which will be ignored(?) Thanks,[/quote]Removing any extra entries should be automated &amp; part of your model code.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "18651": "",
    "18654": "",
    "18663": "",
    "18689": "",
    "18825": "",
    "18938": "",
    "19112": "",
    "19963": "",
    "19965": ""
  },
  "source": "meta"
}