{
  "id": 4185,
  "title": "sanity check in two steps competitions",
  "url": "/competitions/flight/discussion/4185",
  "author_name": "",
  "post_date": "2013-04-03T15:32:05.590Z",
  "votes": null,
  "comment_count": 11,
  "views": 11788,
  "content": "<p>Hi,&nbsp;</p>\r\n<p>I waited for the results but now I am just disappointed.&nbsp;</p>\r\n<p>I hope in the future we will have a bit more detailed response in two step competitions when uploading our final submission.&nbsp;</p>\r\n<p>We obviously made serious mistakes when we missed the fact that we have insane (~55 374 602) prediction for 10 records.&nbsp;</p>\r\n<p>If we just replace these 10 values with the mean prediction (1254-1261) we get close to the Estimated Arrival Time Benchmark.</p>\r\n<p>We worked a &quot;bit&quot; more to end atthe last place :)</p>",
  "messages": [
    {
      "id": "22075",
      "postDate": "04/03/2013 15:32:05",
      "content": "<p>Hi,&nbsp;</p>\r\n<p>I waited for the results but now I am just disappointed.&nbsp;</p>\r\n<p>I hope in the future we will have a bit more detailed response in two step competitions when uploading our final submission.&nbsp;</p>\r\n<p>We obviously made serious mistakes when we missed the fact that we have insane (~55 374 602) prediction for 10 records.&nbsp;</p>\r\n<p>If we just replace these 10 values with the mean prediction (1254-1261) we get close to the Estimated Arrival Time Benchmark.</p>\r\n<p>We worked a &quot;bit&quot; more to end atthe last place :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22077",
      "postDate": "04/03/2013 15:49:57",
      "content": "<p>It seems that the score for people who did not submit in Phase 2 should be raised from 9999 to 999,999,999,999 so that everyone who submitted in Phase 2 ranks higher than people who did not.&nbsp;</p>\r\n<p>Vlado Boza had a great idea in the Job Salary competition - comparing one's Phase 2 submission against a Phase 2 benchmark results dataset, and finding the error.&nbsp; There was such a dataset in Jab Salary, but unfortunately in Flight Quest, there was not.&nbsp;\r\n So that would be a possible solution going forward: for Kaggle to release a benchmark results dataset of the Phase 2 data for people to do a sanity check against.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22078",
      "postDate": "04/03/2013 16:00:05",
      "content": "<p>[quote=BreakfastPirate;22077]</p>\r\n<p>It seems that the score for people who did not submit in Phase 2 should be raised from 9999 to 999,999,999,999 so that everyone who submitted in Phase 2 ranks higher than people who did not.&nbsp;</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>Hey, I would lose my high score! ;)&nbsp; Apart from that, I like the idea.</p>\r\n<p>(Actually, I got this high score because the system didn't parse my solution correctly. Suddenly it seemed to expect other columns / column headers than specified in the submission instructions. Did anyone else experience that sort of problem?)</p>\r\n<p>&nbsp;</p>\r\n<p>updated: so much for my high score - it has been corrected already. (It's still depressing, though.)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22081",
      "postDate": "04/03/2013 16:12:56",
      "content": "<p>[quote=beluga;22075]</p>\r\n<p>Hi,&nbsp;</p>\r\n<p>I waited for the results but now I am just disappointed.&nbsp;</p>\r\n<p>I hope in the future we will have a bit more detailed response in two step competitions when uploading our final submission.&nbsp;</p>\r\n<p>We obviously made serious mistakes when we missed the fact that we have insane (~55 374 602) prediction for 10 records.&nbsp;</p>\r\n<p>If we just replace these 10 values with the mean prediction (1254-1261) we get close to the Estimated Arrival Time Benchmark.</p>\r\n<p>We worked a &quot;bit&quot; more to end atthe last place :)</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>@beluga: I think you should ask admin about your ranking. Most people got the 43th /179.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22124",
      "postDate": "04/04/2013 04:54:08",
      "content": "<p>It's fairly uninspiring that giving up on a competition can earn a top 25% badge.</p>\r\n<p>Also, I'm puzzled by the scores achieved by those who did make a final submission. Why are they so different to the public leaderboard scores?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22136",
      "postDate": "04/04/2013 05:58:41",
      "content": "<p>[quote=David Knox;22124]</p>\r\n<p><span>Also, I'm puzzled by the scores achieved by those who did make a final submission. Why are they so different to the public leaderboard scores?</span></p>\r\n<p><span>[/quote]</span></p>\r\n<p><span>The weather was generally worse in February (final evaluation set) than in December (public leaderboard set). There were snowstorms (as seen in the news), so much more flights got delayed. So the predictions were worse on average, as it's harder to\r\n predict delayed flights.</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22138",
      "postDate": "04/04/2013 07:27:06",
      "content": "<p>I agree that dataset shift is one possible reason for the differences. The reason for such a shift might be two-fold:</p>\r\n<p>&nbsp; a) it could be a shift in the covariates (i.e. features), as claimed by charango; think, P(X) changed between train and test.</p>\r\n<p>&nbsp; b) it could be a shift in the labeling function (i.e., arrival times); think, P(Y | X) changed between train and test</p>\r\n<p>In order to pin-point if a) is the source of the error we could use density estimation techniques or (easier) generate a classification problem (re-label our examples) where we try to classify if a flight stems from the test set or the training set. If there\r\n is a covariate shift the error should be significantly lower than random guessing [1].</p>\r\n<p>Its impossible to detect b) without having access to the labels of the test set (I don't want to query flightstats for each of the 25K test examples :-) ).</p>\r\n<p>Personally, I think the reason for the large differences between public and private scores is the fact that model selection for this competition was shaky&nbsp;due to the fact that we didn't know which flights are eligable for the test set. I assume most participants\r\n based their final model selection on the public leaderboard data rather than internal held-out data; I certainly did - having failed to correlate internal held-out scores w/ public LB scores.</p>\r\n<p>We all know that the error function in this competition is highly sensitive to outliers thus individual flights have a huge impact on the outcome of the competition - a slight change to the selection mechanism might render the public LB not effective for\r\n model selection.</p>\r\n<p>Consider the following example: we are predicting the arrival of <strong>active</strong>&nbsp;<span style=\"font-size:14px; line-height:1.4em\">flights; IMHO its highly unlikely that a flight that's in the air will have a delay of more than say 5 hours (talking\r\n about runway arrival in particular); so its a sensible thing to clip predicted delays of more than 5 hours, however, if there is a flight in the test set with a delay of say 10 hours, the squared residual of that flight (300**2) will be so large that it will\r\n dominate all other residuals and in fact, this one flight among the 25K in the test set will break your neck and most likely decide the result of the competition. I just would like you to know that I found a flight with more than 16h of delay in the test set\r\n which I reported to Ben (if you ask yourself how: using flightstats.com).</span></p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">BTW: Simply plotting an empirical cumulative distribution function of the squared residuals per flight might clarify if this was indeed an issue.<br>\r\n</span></p>\r\n<p>It might be interesting indeed to perform some post-mortem analysis - especially, which flights / days / airports decided the competition, however, since nobody of the organizers answered my question regarding making the final data available I assume all\r\n we get is a single number.</p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">[1]&nbsp;<a href=\"http://blog.smola.org/post/4110255196/real-simple-covariate-shift-correction\">http://blog.smola.org/post/4110255196/real-simple-covariate-shift-correction</a>&nbsp;;&nbsp;<span>Bickel, S., Brückner, M.,\r\n Scheffer, T., 2009. Discriminative learning under covariate shift. J. Mach. Learn. Res. 10, 2137-2155.</span></span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22140",
      "postDate": "04/04/2013 08:14:01",
      "content": "<p>I've ran my model on all of the available data before the final submission. It showed the worst results on the AugmentedTestSet1 (Christmas holidays), the score was comparable to what we have on the final leaderboard.</p>\r\n<p>Basically, when calculating the score for the days separately, my model had the results ranging from 3.0 to 10.0 on average (for a day's data). The score meaning the error (in minutes) for estimated runway arrival time. The predictions for the days with\r\n good weather, i.e. no delays with flight plan estimates, were very good. On the busy days, with bad weather, they were the worst.</p>\r\n<p>I've also spotted some cases of flights arriving with 10&#43; hours delay. Those are easily explainable though - they are diverted flights, but not marked by FlightStats as such. If you'll check a few such samples, you will be able to see 2 ARA (Actual Runway\r\n Arrival) events, several hours apart, and the ASDI (GPS) position implies landing in a different airport (not the destination one). This is what a diverted flight is. However, there is no Status=D line in the flight events file, so the flight does not get\r\n filtered out.</p>\r\n<p>There are much more issues with the provided data, as most of the competitors have surely noticed while working with it.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22142",
      "postDate": "04/04/2013 08:49:07",
      "content": "<p>Before the code freeze deadline I did my checks on the last two weeks of the available data; I used Ben's code to create the test set - the predictions were way off (benchmarks got worse by ~50% - EAT from 9.7 to 13.7, SAT from 29 up to 45 ). I noticed that\r\n if I only used the most recent training data instead of all data my predictions got better (relative to the benchmark). However, then I looked into the residual distribution: most was caused by a tiny fraction of flights (~20 of 20K); I removed them (ie. I\r\n only retained flights with status A); benchmark scores got way better (EAT was now 10.27) and predictions were similar to what I got on the LB; now training on all data was better than only on most recent.</p>\r\n<p>I was a bit confused and hence I decided to be risk averse and went with the same solution as for the public leaderboard.</p>\r\n<p>It seems that I should have spent more time on data cleansing.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22144",
      "postDate": "04/04/2013 09:28:06",
      "content": "<p>Yes, the errors are higher.</p>\r\n<p>More importantly IMHO, the public leaderboard was crippled by the fact that the data it was based on came before ~25th of December. Before that date a huge part of plane positions were missing.</p>\r\n<p><span style=\"line-height:1.4em\">The public leaderboard is a source of motivation, I think it's important that it provides more reliable/unbiased feedback.</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22301",
      "postDate": "04/05/2013 22:12:51",
      "content": "<p>[quote=beluga;22075]</p>\r\n<p>Hi,&nbsp;</p>\r\n<p>I waited for the results but now I am just disappointed.&nbsp;</p>\r\n<p>I hope in the future we will have a bit more detailed response in two step competitions when uploading our final submission.&nbsp;</p>\r\n<p>We obviously made serious mistakes when we missed the fact that we have insane (~55 374 602) prediction for 10 records.&nbsp;</p>\r\n<p>If we just replace these 10 values with the mean prediction (1254-1261) we get close to the Estimated Arrival Time Benchmark.</p>\r\n<p>We worked a &quot;bit&quot; more to end atthe last place :)</p>\r\n<p>[/quote]Sorry about this - I'd thought 9999.0 would be sufficient to put it behind all final submissions. Clearly it wasn't, and I've updated it accordingly. We'll still working out how to do the rankings / support mutli-stage competitions properly, and\r\n these may be updated retroactively to reflect this as well. (One possibilities is to rank based on final evaluation performance, and then based on public leaderboard performance if they made no final evaluation submissions, with everyone making a final evaluation\r\n submission ranking above everyone who only made a public leaderboard submission).</p>\r\n<p>As far as the other issue, there's not a good way around it in this specific case - a very small number of samples (say 1-20) being very far off can destroy your performance regadless of whether it's a single or multi stage competition. (If those outliers\r\n randomly fell in the private leaderboard set, you wouldn't be aware of it in a single stage competition prior to the end either). One possibility is providing a binary feedback for multistage competitions &quot;better than best benchmark&quot; or &quot;worse than best benchmark&quot;.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23686",
      "postDate": "04/30/2013 07:10:22",
      "content": "<p>@<span>charango I had a look into the dataset shift issue - in particular I wanted to know if indeed the weather was the main source of the increased RMSE (compared to the leaderboard).</span></p>\r\n<p>In order to test if there was a difference between the data distributions I divided all data (train and test) into two week periods; for each consecutive two periods I trained a classifier to separate the training from the test instances; if the data distribution\r\n did not change between the two periods the error rate should be high (~ random guessing) otherwise this might indicate a (consistent) change in the feature values. Since the hypothesis was that the weather was the main reason for the distribution shift I only\r\n focused on the METAR features.</p>\r\n<p>Below you can see the results; S... source domain (~ 2 weeks), T... target domain (~ 2 weeks after S). The score to the right of S | T is the training error - lower means that the two domains are easily separable. Each score is the average of 10 repetitions\r\n of sampling 1000 samples from both S and T at random (the number in the brackets is the std). I used a gr<span style=\"font-size:14px; line-height:1.4em\">adient boosting classifier w/ log loss; below each line you can see the 5 most important features.</span></p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">&lt;pre&gt;</span></p>\r\n<pre>S=(2012-11-12, 2012-11-28] | T=(2012-11-28, 2012-12-12]: 0.7399 (0.0096)<br>   ['metar_sc_2', 'metar_wind_direction', 'metar_altimeter', 'metar_sea_level_pressure', 'metar_dewpoint']<br>S=(2012-11-28, 2012-12-12] | T=(2012-12-12, 2012-12-26]: 0.7418 (0.0087)<br>   ['metar_wind_direction', 'metar_sc_2', 'metar_temperature', 'metar_dewpoint', 'metar_altimeter']<br>S=(2012-12-12, 2012-12-26] | T=(2012-12-26, 2013-01-09]: 0.6996 (0.0172)<br>   ['metar_sc_1', 'metar_temperature', 'metar_wind_direction', 'metar_dewpoint', 'metar_altimeter']<br>S=(2012-12-26, 2013-01-09] | T=(2013-01-09, 2013-01-23]: 0.7720 (0.0083)<br>   ['metar_sc_3', 'metar_wind_direction', 'metar_altimeter', 'metar_temperature', 'metar_dewpoint']<br>S=(2013-01-09, 2013-01-23] | T=(2013-01-23, 2013-02-06]: 0.8433 (0.0069)<br>   ['metar_wind_direction', 'metar_sea_level_pressure', 'metar_temperature', 'metar_altimeter', 'metar_dewpoint']<br>S=(2013-01-23, 2013-02-06] | T=(2013-02-06, 2013-02-28]: 0.7398 (0.0073)<br>   ['metar_sea_level_pressure', 'metar_wind_direction', 'metar_temperature', 'metar_altimeter', 'metar_dewpoint']</pre>\r\n<pre>&lt;/pre&gt;</pre>\r\n<pre>The absolute scores are not really important since we are only interested if there was a change relative to the public leaderboard (or our internal validation sets).</pre>\r\n<pre>According to the above analysis this was not the case.&nbsp;</pre>\r\n<pre><span style=\"font-size:14px; line-height:1.4em\">I've attached some plots of metar features (dewpoint, temperature, altimeter; test set starts at 15751) -&nbsp;</span><span style=\"font-size:14px; line-height:1.4em\">they support the claim that there was indeed no consistant difference in the METAR features between the training data and the final evalation set.</span></pre>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 22077,
      "author_name": "breakfastpirate",
      "author_url": "",
      "post_date": "04/03/2013 15:49:57",
      "content": "<p>It seems that the score for people who did not submit in Phase 2 should be raised from 9999 to 999,999,999,999 so that everyone who submitted in Phase 2 ranks higher than people who did not.&nbsp;</p>\r\n<p>Vlado Boza had a great idea in the Job Salary competition - comparing one's Phase 2 submission against a Phase 2 benchmark results dataset, and finding the error.&nbsp; There was such a dataset in Jab Salary, but unfortunately in Flight Quest, there was not.&nbsp;\r\n So that would be a possible solution going forward: for Kaggle to release a benchmark results dataset of the Phase 2 data for people to do a sanity check against.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22078,
      "author_name": "ansgar",
      "author_url": "",
      "post_date": "04/03/2013 16:00:05",
      "content": "<p>[quote=BreakfastPirate;22077]</p>\r\n<p>It seems that the score for people who did not submit in Phase 2 should be raised from 9999 to 999,999,999,999 so that everyone who submitted in Phase 2 ranks higher than people who did not.&nbsp;</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>Hey, I would lose my high score! ;)&nbsp; Apart from that, I like the idea.</p>\r\n<p>(Actually, I got this high score because the system didn't parse my solution correctly. Suddenly it seemed to expect other columns / column headers than specified in the submission instructions. Did anyone else experience that sort of problem?)</p>\r\n<p>&nbsp;</p>\r\n<p>updated: so much for my high score - it has been corrected already. (It's still depressing, though.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22081,
      "author_name": "songgc",
      "author_url": "",
      "post_date": "04/03/2013 16:12:56",
      "content": "<p>[quote=beluga;22075]</p>\r\n<p>Hi,&nbsp;</p>\r\n<p>I waited for the results but now I am just disappointed.&nbsp;</p>\r\n<p>I hope in the future we will have a bit more detailed response in two step competitions when uploading our final submission.&nbsp;</p>\r\n<p>We obviously made serious mistakes when we missed the fact that we have insane (~55 374 602) prediction for 10 records.&nbsp;</p>\r\n<p>If we just replace these 10 values with the mean prediction (1254-1261) we get close to the Estimated Arrival Time Benchmark.</p>\r\n<p>We worked a &quot;bit&quot; more to end atthe last place :)</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>@beluga: I think you should ask admin about your ranking. Most people got the 43th /179.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22124,
      "author_name": "davidknox",
      "author_url": "",
      "post_date": "04/04/2013 04:54:08",
      "content": "<p>It's fairly uninspiring that giving up on a competition can earn a top 25% badge.</p>\r\n<p>Also, I'm puzzled by the scores achieved by those who did make a final submission. Why are they so different to the public leaderboard scores?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22136,
      "author_name": "sergeykozub",
      "author_url": "",
      "post_date": "04/04/2013 05:58:41",
      "content": "<p>[quote=David Knox;22124]</p>\r\n<p><span>Also, I'm puzzled by the scores achieved by those who did make a final submission. Why are they so different to the public leaderboard scores?</span></p>\r\n<p><span>[/quote]</span></p>\r\n<p><span>The weather was generally worse in February (final evaluation set) than in December (public leaderboard set). There were snowstorms (as seen in the news), so much more flights got delayed. So the predictions were worse on average, as it's harder to\r\n predict delayed flights.</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22138,
      "author_name": "pprett",
      "author_url": "",
      "post_date": "04/04/2013 07:27:06",
      "content": "<p>I agree that dataset shift is one possible reason for the differences. The reason for such a shift might be two-fold:</p>\r\n<p>&nbsp; a) it could be a shift in the covariates (i.e. features), as claimed by charango; think, P(X) changed between train and test.</p>\r\n<p>&nbsp; b) it could be a shift in the labeling function (i.e., arrival times); think, P(Y | X) changed between train and test</p>\r\n<p>In order to pin-point if a) is the source of the error we could use density estimation techniques or (easier) generate a classification problem (re-label our examples) where we try to classify if a flight stems from the test set or the training set. If there\r\n is a covariate shift the error should be significantly lower than random guessing [1].</p>\r\n<p>Its impossible to detect b) without having access to the labels of the test set (I don't want to query flightstats for each of the 25K test examples :-) ).</p>\r\n<p>Personally, I think the reason for the large differences between public and private scores is the fact that model selection for this competition was shaky&nbsp;due to the fact that we didn't know which flights are eligable for the test set. I assume most participants\r\n based their final model selection on the public leaderboard data rather than internal held-out data; I certainly did - having failed to correlate internal held-out scores w/ public LB scores.</p>\r\n<p>We all know that the error function in this competition is highly sensitive to outliers thus individual flights have a huge impact on the outcome of the competition - a slight change to the selection mechanism might render the public LB not effective for\r\n model selection.</p>\r\n<p>Consider the following example: we are predicting the arrival of <strong>active</strong>&nbsp;<span style=\"font-size:14px; line-height:1.4em\">flights; IMHO its highly unlikely that a flight that's in the air will have a delay of more than say 5 hours (talking\r\n about runway arrival in particular); so its a sensible thing to clip predicted delays of more than 5 hours, however, if there is a flight in the test set with a delay of say 10 hours, the squared residual of that flight (300**2) will be so large that it will\r\n dominate all other residuals and in fact, this one flight among the 25K in the test set will break your neck and most likely decide the result of the competition. I just would like you to know that I found a flight with more than 16h of delay in the test set\r\n which I reported to Ben (if you ask yourself how: using flightstats.com).</span></p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">BTW: Simply plotting an empirical cumulative distribution function of the squared residuals per flight might clarify if this was indeed an issue.<br>\r\n</span></p>\r\n<p>It might be interesting indeed to perform some post-mortem analysis - especially, which flights / days / airports decided the competition, however, since nobody of the organizers answered my question regarding making the final data available I assume all\r\n we get is a single number.</p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">[1]&nbsp;<a href=\"http://blog.smola.org/post/4110255196/real-simple-covariate-shift-correction\">http://blog.smola.org/post/4110255196/real-simple-covariate-shift-correction</a>&nbsp;;&nbsp;<span>Bickel, S., Brückner, M.,\r\n Scheffer, T., 2009. Discriminative learning under covariate shift. J. Mach. Learn. Res. 10, 2137-2155.</span></span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22140,
      "author_name": "sergeykozub",
      "author_url": "",
      "post_date": "04/04/2013 08:14:01",
      "content": "<p>I've ran my model on all of the available data before the final submission. It showed the worst results on the AugmentedTestSet1 (Christmas holidays), the score was comparable to what we have on the final leaderboard.</p>\r\n<p>Basically, when calculating the score for the days separately, my model had the results ranging from 3.0 to 10.0 on average (for a day's data). The score meaning the error (in minutes) for estimated runway arrival time. The predictions for the days with\r\n good weather, i.e. no delays with flight plan estimates, were very good. On the busy days, with bad weather, they were the worst.</p>\r\n<p>I've also spotted some cases of flights arriving with 10&#43; hours delay. Those are easily explainable though - they are diverted flights, but not marked by FlightStats as such. If you'll check a few such samples, you will be able to see 2 ARA (Actual Runway\r\n Arrival) events, several hours apart, and the ASDI (GPS) position implies landing in a different airport (not the destination one). This is what a diverted flight is. However, there is no Status=D line in the flight events file, so the flight does not get\r\n filtered out.</p>\r\n<p>There are much more issues with the provided data, as most of the competitors have surely noticed while working with it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22142,
      "author_name": "pprett",
      "author_url": "",
      "post_date": "04/04/2013 08:49:07",
      "content": "<p>Before the code freeze deadline I did my checks on the last two weeks of the available data; I used Ben's code to create the test set - the predictions were way off (benchmarks got worse by ~50% - EAT from 9.7 to 13.7, SAT from 29 up to 45 ). I noticed that\r\n if I only used the most recent training data instead of all data my predictions got better (relative to the benchmark). However, then I looked into the residual distribution: most was caused by a tiny fraction of flights (~20 of 20K); I removed them (ie. I\r\n only retained flights with status A); benchmark scores got way better (EAT was now 10.27) and predictions were similar to what I got on the LB; now training on all data was better than only on most recent.</p>\r\n<p>I was a bit confused and hence I decided to be risk averse and went with the same solution as for the public leaderboard.</p>\r\n<p>It seems that I should have spent more time on data cleansing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22144,
      "author_name": "melisgl",
      "author_url": "",
      "post_date": "04/04/2013 09:28:06",
      "content": "<p>Yes, the errors are higher.</p>\r\n<p>More importantly IMHO, the public leaderboard was crippled by the fact that the data it was based on came before ~25th of December. Before that date a huge part of plane positions were missing.</p>\r\n<p><span style=\"line-height:1.4em\">The public leaderboard is a source of motivation, I think it's important that it provides more reliable/unbiased feedback.</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22301,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "04/05/2013 22:12:51",
      "content": "<p>[quote=beluga;22075]</p>\r\n<p>Hi,&nbsp;</p>\r\n<p>I waited for the results but now I am just disappointed.&nbsp;</p>\r\n<p>I hope in the future we will have a bit more detailed response in two step competitions when uploading our final submission.&nbsp;</p>\r\n<p>We obviously made serious mistakes when we missed the fact that we have insane (~55 374 602) prediction for 10 records.&nbsp;</p>\r\n<p>If we just replace these 10 values with the mean prediction (1254-1261) we get close to the Estimated Arrival Time Benchmark.</p>\r\n<p>We worked a &quot;bit&quot; more to end atthe last place :)</p>\r\n<p>[/quote]Sorry about this - I'd thought 9999.0 would be sufficient to put it behind all final submissions. Clearly it wasn't, and I've updated it accordingly. We'll still working out how to do the rankings / support mutli-stage competitions properly, and\r\n these may be updated retroactively to reflect this as well. (One possibilities is to rank based on final evaluation performance, and then based on public leaderboard performance if they made no final evaluation submissions, with everyone making a final evaluation\r\n submission ranking above everyone who only made a public leaderboard submission).</p>\r\n<p>As far as the other issue, there's not a good way around it in this specific case - a very small number of samples (say 1-20) being very far off can destroy your performance regadless of whether it's a single or multi stage competition. (If those outliers\r\n randomly fell in the private leaderboard set, you wouldn't be aware of it in a single stage competition prior to the end either). One possibility is providing a binary feedback for multistage competitions &quot;better than best benchmark&quot; or &quot;worse than best benchmark&quot;.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23686,
      "author_name": "pprett",
      "author_url": "",
      "post_date": "04/30/2013 07:10:22",
      "content": "<p>@<span>charango I had a look into the dataset shift issue - in particular I wanted to know if indeed the weather was the main source of the increased RMSE (compared to the leaderboard).</span></p>\r\n<p>In order to test if there was a difference between the data distributions I divided all data (train and test) into two week periods; for each consecutive two periods I trained a classifier to separate the training from the test instances; if the data distribution\r\n did not change between the two periods the error rate should be high (~ random guessing) otherwise this might indicate a (consistent) change in the feature values. Since the hypothesis was that the weather was the main reason for the distribution shift I only\r\n focused on the METAR features.</p>\r\n<p>Below you can see the results; S... source domain (~ 2 weeks), T... target domain (~ 2 weeks after S). The score to the right of S | T is the training error - lower means that the two domains are easily separable. Each score is the average of 10 repetitions\r\n of sampling 1000 samples from both S and T at random (the number in the brackets is the std). I used a gr<span style=\"font-size:14px; line-height:1.4em\">adient boosting classifier w/ log loss; below each line you can see the 5 most important features.</span></p>\r\n<p><span style=\"font-size:14px; line-height:1.4em\">&lt;pre&gt;</span></p>\r\n<pre>S=(2012-11-12, 2012-11-28] | T=(2012-11-28, 2012-12-12]: 0.7399 (0.0096)<br>   ['metar_sc_2', 'metar_wind_direction', 'metar_altimeter', 'metar_sea_level_pressure', 'metar_dewpoint']<br>S=(2012-11-28, 2012-12-12] | T=(2012-12-12, 2012-12-26]: 0.7418 (0.0087)<br>   ['metar_wind_direction', 'metar_sc_2', 'metar_temperature', 'metar_dewpoint', 'metar_altimeter']<br>S=(2012-12-12, 2012-12-26] | T=(2012-12-26, 2013-01-09]: 0.6996 (0.0172)<br>   ['metar_sc_1', 'metar_temperature', 'metar_wind_direction', 'metar_dewpoint', 'metar_altimeter']<br>S=(2012-12-26, 2013-01-09] | T=(2013-01-09, 2013-01-23]: 0.7720 (0.0083)<br>   ['metar_sc_3', 'metar_wind_direction', 'metar_altimeter', 'metar_temperature', 'metar_dewpoint']<br>S=(2013-01-09, 2013-01-23] | T=(2013-01-23, 2013-02-06]: 0.8433 (0.0069)<br>   ['metar_wind_direction', 'metar_sea_level_pressure', 'metar_temperature', 'metar_altimeter', 'metar_dewpoint']<br>S=(2013-01-23, 2013-02-06] | T=(2013-02-06, 2013-02-28]: 0.7398 (0.0073)<br>   ['metar_sea_level_pressure', 'metar_wind_direction', 'metar_temperature', 'metar_altimeter', 'metar_dewpoint']</pre>\r\n<pre>&lt;/pre&gt;</pre>\r\n<pre>The absolute scores are not really important since we are only interested if there was a change relative to the public leaderboard (or our internal validation sets).</pre>\r\n<pre>According to the above analysis this was not the case.&nbsp;</pre>\r\n<pre><span style=\"font-size:14px; line-height:1.4em\">I've attached some plots of metar features (dewpoint, temperature, altimeter; test set starts at 15751) -&nbsp;</span><span style=\"font-size:14px; line-height:1.4em\">they support the claim that there was indeed no consistant difference in the METAR features between the training data and the final evalation set.</span></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "22075": "",
    "22077": "",
    "22078": "",
    "22081": "",
    "22124": "",
    "22136": "",
    "22138": "",
    "22140": "",
    "22142": "",
    "22144": "",
    "22301": "",
    "23686": ""
  },
  "source": "meta"
}