{
  "id": 5358,
  "title": "Warning on weather training",
  "url": "/competitions/flight2-milestone/discussion/5358",
  "author_name": "",
  "post_date": "2013-08-09T13:54:55.213Z",
  "votes": 2,
  "comment_count": 3,
  "views": 2398,
  "content": "<p>As a warning to anyone planning on training with weather features, note that the training sets take place in midsummer but the final evaluation will be based on winter data.&nbsp; Experience shows that weather patterns in the US are very different in those different seasons, especially in relation to some conditions that may appear in the test set but not the training set (e. g. icing conditions at a large number of airports).&nbsp; There are also different density and structure of holidays, different passenger cost functions (which the simulator may or may not reflect; it's not released yet), and other important distinguishing factors.</p>\n<p>The leaderboards are based on different seasons yet (summer &amp; early autumn), so there is not really any way to know how one's algorithm is going to perform on the final evaluation set.&nbsp; The agent training leaderboard also covers Sept. 11, though I'm not sure if the simulator will reflect any flukes in US aviation on that date (e. g. cost function for routes changed to be closer to major urban centers, especially at lower altitudes, possibly being higher than on other dates).&nbsp;</p>\n<p><br>I'd look for some pretty big changes between leaderboard scores and final evaluation scores.&nbsp; #100 on the leaderboard at the close of competition...you could win!<br><br>Personally, I think it would be better to have training data on similar conditions as the test data, either having both in the midsummer (maybe August, without holidays) or both as year-round samples.&nbsp; <br><br>As this competition is not set up like that, unless the admins change it, you will either have to rely on external weather sources for other time periods (though you don't have matching flight data to pass through the simulator) or simply not include a lot of detail on weather-related features, since they're likely to overfit on the summer data (match the details of summer much better than the test data, which is in winter).&nbsp; <br><br></p>",
  "messages": [
    {
      "id": "28481",
      "postDate": "08/09/2013 13:54:55",
      "content": "<p>As a warning to anyone planning on training with weather features, note that the training sets take place in midsummer but the final evaluation will be based on winter data.&nbsp; Experience shows that weather patterns in the US are very different in those different seasons, especially in relation to some conditions that may appear in the test set but not the training set (e. g. icing conditions at a large number of airports).&nbsp; There are also different density and structure of holidays, different passenger cost functions (which the simulator may or may not reflect; it's not released yet), and other important distinguishing factors.</p>\n<p>The leaderboards are based on different seasons yet (summer &amp; early autumn), so there is not really any way to know how one's algorithm is going to perform on the final evaluation set.&nbsp; The agent training leaderboard also covers Sept. 11, though I'm not sure if the simulator will reflect any flukes in US aviation on that date (e. g. cost function for routes changed to be closer to major urban centers, especially at lower altitudes, possibly being higher than on other dates).&nbsp;</p>\n<p><br>I'd look for some pretty big changes between leaderboard scores and final evaluation scores.&nbsp; #100 on the leaderboard at the close of competition...you could win!<br><br>Personally, I think it would be better to have training data on similar conditions as the test data, either having both in the midsummer (maybe August, without holidays) or both as year-round samples.&nbsp; <br><br>As this competition is not set up like that, unless the admins change it, you will either have to rely on external weather sources for other time periods (though you don't have matching flight data to pass through the simulator) or simply not include a lot of detail on weather-related features, since they're likely to overfit on the summer data (match the details of summer much better than the test data, which is in winter).&nbsp; <br><br></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28576",
      "postDate": "08/12/2013 08:18:14",
      "content": "<p>Aircraft schedules in winter and summer are very different not only due to weather but due to passenger preferences and demand (people on summer vacation etc...). &nbsp;Congestion related delays in the real world are very seasonal at many airports.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28712",
      "postDate": "08/15/2013 00:14:36",
      "content": "<p>I don't think it will matter much to final ranking. These kind of errors are systematic - they should touch everyone to the same extent.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28733",
      "postDate": "08/15/2013 13:37:01",
      "content": "<p>If everybody had the same algorithm, that would be the case. <br>These kinds of errors will affect everyone differently, depending on the degree to which the training and test features *that are prominent in the model* differ between training and test.&nbsp; <br>Here, I recommend that people pick features that are not as likely to have big differences between training and test, but between my post and Noam's, there isn't a whole lot left, and we should just expect big swings and reshufflings among ranked performance order.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 28576,
      "author_name": "noamtene",
      "author_url": "",
      "post_date": "08/12/2013 08:18:14",
      "content": "<p>Aircraft schedules in winter and summer are very different not only due to weather but due to passenger preferences and demand (people on summer vacation etc...). &nbsp;Congestion related delays in the real world are very seasonal at many airports.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28712,
      "author_name": "paweljankiewicz",
      "author_url": "",
      "post_date": "08/15/2013 00:14:36",
      "content": "<p>I don't think it will matter much to final ranking. These kind of errors are systematic - they should touch everyone to the same extent.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28733,
      "author_name": "wbtthefrog",
      "author_url": "",
      "post_date": "08/15/2013 13:37:01",
      "content": "<p>If everybody had the same algorithm, that would be the case. <br>These kinds of errors will affect everyone differently, depending on the degree to which the training and test features *that are prominent in the model* differ between training and test.&nbsp; <br>Here, I recommend that people pick features that are not as likely to have big differences between training and test, but between my post and Noam's, there isn't a whole lot left, and we should just expect big swings and reshufflings among ranked performance order.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "28481": "",
    "28576": "",
    "28712": "",
    "28733": ""
  },
  "source": "meta"
}