{
  "id": 6735,
  "title": "unavailabiliyt of groundConditions/actualLandings/actualTakeoffs files",
  "url": "/competitions/flight2-main/discussion/6735",
  "author_name": "",
  "post_date": "2014-01-01T23:45:09.963Z",
  "votes": null,
  "comment_count": 5,
  "views": 2602,
  "content": "<p>Hi Joycenv,</p>\n<p>Do you know where to get&nbsp;groundConditions/actualLandings/actualTakeoffs files for all the testing dates? I only found these 3 files for Sep.10. &nbsp; Currently, I tentatively use Sep.11 files to run simulation for all dates. However, i don't think this is right way to do as these 3 files might be different for different dates.</p>\n<p>Thanks,</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
  "messages": [
    {
      "id": "36899",
      "postDate": "01/01/2014 23:45:09",
      "content": "<p>Hi Joycenv,</p>\n<p>Do you know where to get&nbsp;groundConditions/actualLandings/actualTakeoffs files for all the testing dates? I only found these 3 files for Sep.10. &nbsp; Currently, I tentatively use Sep.11 files to run simulation for all dates. However, i don't think this is right way to do as these 3 files might be different for different dates.</p>\n<p>Thanks,</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36961",
      "postDate": "01/03/2014 20:12:40",
      "content": "<p>[quote=orchid;36899]</p>\n<p>Currently, I tentatively use Sep.11 files to run simulation for all dates.</p>\n<p>[/quote]</p>\n<p>That is how I started out and believe it or not, that is what I still do. In theory, you should be able to estimate takeoff and landing frequency etc. from training2_flighthistory.csv. Possibly by filtering based on scheduled_runway_departure and scheduled_runway_arrival columns. I just haven't been able to get predictions that work significantly better than just using the actual schedule from the September 10th files. If you would like to give it a try, the <a href=\"https://github.com/benhamner/GEFlightQuest/blob/master/PythonModule/geflight/postgres_ingest/schema.sql\">SQL code</a> posted earlier might be a good place to start.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36965",
      "postDate": "01/03/2014 20:42:02",
      "content": "<p>I also used Sep. 10 files. However, in one of my simulation testings, I found a huge difference between my local simulation score and my Leaderboard score. My local simulation scored 13093 while my Leaderboard score was 14910, &nbsp;which made me very concerned on quality of local simulation and also very concerned &nbsp;on validity of extending sep. 10 files to other dates.&nbsp;</p>\n<p>How big difference you see between your local simulation and your leaderboard scores ?</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36969",
      "postDate": "01/03/2014 21:14:54",
      "content": "<p>That is a huge difference between your local and leaderboard scores. Maybe you can try throwing a different weather file at the flights in your local simulation. You will very likely find that some of the flights are crashing with changed wind conditions. This of course, extends to the leaderboard where actual (read different) weather is used in the simulation.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36973",
      "postDate": "01/03/2014 22:17:02",
      "content": "<p>not very sure what you meant.&nbsp;</p>\n<p>We used turbulent_zones_xxxxxxxx_xxxx.csv and weather_xxxxxxxx_xxxx.txt.gz files all provided by Kaggle. The only weather-related file that we had to borrow from Sep. 10 to other dates &nbsp;is groundConditions_xxxxxxxx_xxxx.csv.</p>\n<p>&nbsp;</p>\n<p>which weather file you meant that will cause difference between local simulation and leaderboard ?</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "36986",
      "postDate": "01/04/2014 07:22:11",
      "content": "<p>The weather_xxxxxxxx_xxxx.txt.gz files for the test dates contain a single snapshot. Kaggle gets the leaderboard scores by running simulations with weather files that have 8 snapshots (each snapshot is taken 1hr apart). When you run a local simulation, the simulator assumes that the wind velocities do not change along the time axis. This assumption does not hold for the leaderboard.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 36961,
      "author_name": "anlthms",
      "author_url": "",
      "post_date": "01/03/2014 20:12:40",
      "content": "<p>[quote=orchid;36899]</p>\n<p>Currently, I tentatively use Sep.11 files to run simulation for all dates.</p>\n<p>[/quote]</p>\n<p>That is how I started out and believe it or not, that is what I still do. In theory, you should be able to estimate takeoff and landing frequency etc. from training2_flighthistory.csv. Possibly by filtering based on scheduled_runway_departure and scheduled_runway_arrival columns. I just haven't been able to get predictions that work significantly better than just using the actual schedule from the September 10th files. If you would like to give it a try, the <a href=\"https://github.com/benhamner/GEFlightQuest/blob/master/PythonModule/geflight/postgres_ingest/schema.sql\">SQL code</a> posted earlier might be a good place to start.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36965,
      "author_name": "orchid",
      "author_url": "",
      "post_date": "01/03/2014 20:42:02",
      "content": "<p>I also used Sep. 10 files. However, in one of my simulation testings, I found a huge difference between my local simulation score and my Leaderboard score. My local simulation scored 13093 while my Leaderboard score was 14910, &nbsp;which made me very concerned on quality of local simulation and also very concerned &nbsp;on validity of extending sep. 10 files to other dates.&nbsp;</p>\n<p>How big difference you see between your local simulation and your leaderboard scores ?</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36969,
      "author_name": "anlthms",
      "author_url": "",
      "post_date": "01/03/2014 21:14:54",
      "content": "<p>That is a huge difference between your local and leaderboard scores. Maybe you can try throwing a different weather file at the flights in your local simulation. You will very likely find that some of the flights are crashing with changed wind conditions. This of course, extends to the leaderboard where actual (read different) weather is used in the simulation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36973,
      "author_name": "orchid",
      "author_url": "",
      "post_date": "01/03/2014 22:17:02",
      "content": "<p>not very sure what you meant.&nbsp;</p>\n<p>We used turbulent_zones_xxxxxxxx_xxxx.csv and weather_xxxxxxxx_xxxx.txt.gz files all provided by Kaggle. The only weather-related file that we had to borrow from Sep. 10 to other dates &nbsp;is groundConditions_xxxxxxxx_xxxx.csv.</p>\n<p>&nbsp;</p>\n<p>which weather file you meant that will cause difference between local simulation and leaderboard ?</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 36986,
      "author_name": "anlthms",
      "author_url": "",
      "post_date": "01/04/2014 07:22:11",
      "content": "<p>The weather_xxxxxxxx_xxxx.txt.gz files for the test dates contain a single snapshot. Kaggle gets the leaderboard scores by running simulations with weather files that have 8 snapshots (each snapshot is taken 1hr apart). When you run a local simulation, the simulator assumes that the wind velocities do not change along the time axis. This assumption does not hold for the leaderboard.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "36899": "",
    "36961": "",
    "36965": "",
    "36969": "",
    "36973": "",
    "36986": ""
  },
  "source": "meta"
}