{
  "id": 3391,
  "title": "What training data should we use?",
  "url": "/competitions/flight/discussion/3391",
  "author_name": "",
  "post_date": "2012-12-18T04:38:37.943Z",
  "votes": null,
  "comment_count": 3,
  "views": 3279,
  "content": "<p><span style=\"line-height:1.4em\">Referring to the folder 'SampleTestSet/2012_11_19'. The file 'test_flights.csv' contains ~2K rows, for which their 'actual_runway_arrival' and 'actual_gate_arrival' need to be predicted.</span></p>\r\n<p>Also, the file 'flighthistory.csv' in the same folder 'SampleTestSet/2012_11_19/FlightHistory' (same day)&nbsp;contains ~24K rows. However, only ~7K rows have both their 'actual_runway_arrival' and 'actual_gate_arrival' un-hidden. It seems that the 'actual_runway_arrival'/'actual_gate_arrival'\r\n time for those ~7K rows were all before the cut-off time (for that day), which means - I'm assuming - that we can use all of those ~7K rows to train our model.</p>\r\n<p><span style=\"line-height:1.4em\">So my question is: is it alwasy the case where we can use all (un-hidden) rows provided in the 'flighthistory.csv' for the corresponding day for training our model?</span></p>",
  "messages": [
    {
      "id": "18102",
      "postDate": "12/18/2012 04:38:37",
      "content": "<p><span style=\"line-height:1.4em\">Referring to the folder 'SampleTestSet/2012_11_19'. The file 'test_flights.csv' contains ~2K rows, for which their 'actual_runway_arrival' and 'actual_gate_arrival' need to be predicted.</span></p>\r\n<p>Also, the file 'flighthistory.csv' in the same folder 'SampleTestSet/2012_11_19/FlightHistory' (same day)&nbsp;contains ~24K rows. However, only ~7K rows have both their 'actual_runway_arrival' and 'actual_gate_arrival' un-hidden. It seems that the 'actual_runway_arrival'/'actual_gate_arrival'\r\n time for those ~7K rows were all before the cut-off time (for that day), which means - I'm assuming - that we can use all of those ~7K rows to train our model.</p>\r\n<p><span style=\"line-height:1.4em\">So my question is: is it alwasy the case where we can use all (un-hidden) rows provided in the 'flighthistory.csv' for the corresponding day for training our model?</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18123",
      "postDate": "12/18/2012 11:04:25",
      "content": "<p>My take on this is to treat rows with flight_history_id that occur in test_flights as test set, and everything else as training.&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>Edit: a tiny problem with that seems to be that (e.g. for 19.11.2012) there are missing observations in&nbsp;scheduled_gate_departure, scheduled_runway_departure and scheduled_runway_arrival</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18491",
      "postDate": "12/23/2012 18:52:27",
      "content": "<p>Then how can you do cross-validation. Are you building on train and directly submitting to see leaderboard score?</p>\r\n<p>&nbsp;</p>\r\n<p>[quote=Konrad Banachewicz;18123]</p>\r\n<p>My take on this is to treat rows with flight_history_id that occur in test_flights as test set, and everything else as training.&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>Edit: a tiny problem with that seems to be that (e.g. for 19.11.2012) there are missing observations in&nbsp;scheduled_gate_departure, scheduled_runway_departure and scheduled_runway_arrival</p>\r\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18494",
      "postDate": "12/23/2012 19:23:35",
      "content": "<p>No, cross-validating on train and using this to rank-order my models - the actual (absolute) numbers are more optimistic than the leaderboard, I haven't (yet :-) figured out way. I hope the controversy about what we can / can not use WILL be resolved / addressed\r\n by the admins at some point...</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 18123,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "12/18/2012 11:04:25",
      "content": "<p>My take on this is to treat rows with flight_history_id that occur in test_flights as test set, and everything else as training.&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>Edit: a tiny problem with that seems to be that (e.g. for 19.11.2012) there are missing observations in&nbsp;scheduled_gate_departure, scheduled_runway_departure and scheduled_runway_arrival</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18491,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "12/23/2012 18:52:27",
      "content": "<p>Then how can you do cross-validation. Are you building on train and directly submitting to see leaderboard score?</p>\r\n<p>&nbsp;</p>\r\n<p>[quote=Konrad Banachewicz;18123]</p>\r\n<p>My take on this is to treat rows with flight_history_id that occur in test_flights as test set, and everything else as training.&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>Edit: a tiny problem with that seems to be that (e.g. for 19.11.2012) there are missing observations in&nbsp;scheduled_gate_departure, scheduled_runway_departure and scheduled_runway_arrival</p>\r\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 18494,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "12/23/2012 19:23:35",
      "content": "<p>No, cross-validating on train and using this to rank-order my models - the actual (absolute) numbers are more optimistic than the leaderboard, I haven't (yet :-) figured out way. I hope the controversy about what we can / can not use WILL be resolved / addressed\r\n by the admins at some point...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "18102": "",
    "18123": "",
    "18491": "",
    "18494": ""
  },
  "source": "meta"
}