{
  "id": 3465,
  "title": "Generating training and test sets",
  "url": "/competitions/flight/discussion/3465",
  "author_name": "",
  "post_date": "2012-12-24T22:08:43.407Z",
  "votes": null,
  "comment_count": 7,
  "views": 5615,
  "content": "<p>Hey, I don't know what I'm doing wrong, but can someone explain how to create training and testing sets off the public leaderboard data? I've been trying to use that code on Ben Hammer's GitHub repository, but it seems like I have to string a ton of little\r\n hacks together just to get it to produce some data--data that I doubt is correct probably due to one of my hacks that forced it to work.</p>\r\n<p>Thanks,</p>\r\n<p>-Brett</p>",
  "messages": [
    {
      "id": "18556",
      "postDate": "12/24/2012 22:08:43",
      "content": "<p>Hey, I don't know what I'm doing wrong, but can someone explain how to create training and testing sets off the public leaderboard data? I've been trying to use that code on Ben Hammer's GitHub repository, but it seems like I have to string a ton of little\r\n hacks together just to get it to produce some data--data that I doubt is correct probably due to one of my hacks that forced it to work.</p>\r\n<p>Thanks,</p>\r\n<p>-Brett</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "18610",
      "postDate": "12/26/2012 06:41:18",
      "content": "<p>[quote=Brett;18556]</p>\r\n<p>Hey, I don't know what I'm doing wrong, but can someone explain how to create training and testing sets off the public leaderboard data? I've been trying to use that code on Ben Hammer's GitHub repository, but it seems like I have to string a ton of little\r\n hacks together just to get it to produce some data--data that I doubt is correct probably due to one of my hacks that forced it to work.</p>\r\n<p>Thanks,</p>\r\n<p>-Brett</p>\r\n<p>[/quote]Sorry for any confusion. The public leaderboard data is a test set, and you should be making / submitting predictions based on that (not trying to create training sets from this data - you're missing all the information after the cutoff time from\r\n each day).</p>\r\n<p>When time permits, I'll clean up the Github code and release a tutorial on deriving a test set from the training data that you can use for internal model development or validation.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19032",
      "postDate": "01/06/2013 16:54:33",
      "content": "<p>[quote=Ben Hamner;18610]</p>\r\n<p>[quote=Brett;18556]</p>\r\n<p>Hey, I don't know what I'm doing wrong, but can someone explain how to create training and testing sets off the public leaderboard data? I've been trying to use that code on Ben Hammer's GitHub repository, but it seems like I have to string a ton of little\r\n hacks together just to get it to produce some data--data that I doubt is correct probably due to one of my hacks that forced it to work.</p>\r\n<p>Thanks,</p>\r\n<p>-Brett</p>\r\n<p>[/quote]Sorry for any confusion. The public leaderboard data is a test set, and you should be making / submitting predictions based on that (not trying to create training sets from this data - you're missing all the information after the cutoff time from\r\n each day).</p>\r\n<p>When time permits, I'll clean up the Github code and release a tutorial on deriving a test set from the training data that you can use for internal model development or validation.</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>Has this tutorial ever been released?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19035",
      "postDate": "01/06/2013 18:02:43",
      "content": "<p>I think there are 3-5 forum topics all asking variations of the same question. It would be nice if we could get conirmation that we will be given a list of flight_history_ids to predict for...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19036",
      "postDate": "01/06/2013 18:10:39",
      "content": "<p>If you read the Submission Instructions at:</p>\r\n<p><a href=\"http://www.gequest.com/c/flight/details/submission-instructions\">http://www.gequest.com/c/flight/details/submission-instructions</a></p>\r\n<p>you will see that:</p>\r\n<p>“Sample submission files (such as estimated_arrival_benchmark.csv) for the public&nbsp;leaderboard set have been uploaded to the\r\n<a href=\"https://www.gequest.com/c/flight/data\">data page</a>. A sample submission for the final evaluation set will be released with that set.”</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19040",
      "postDate": "01/06/2013 19:55:41",
      "content": "<p>ok, look forward to it</p>\r\n<p>&nbsp;</p>\r\n<p>[quote=Ben Hamner;18610]</p>\r\n<p>[quote=Brett;18556]</p>\r\n<p>Hey, I don't know what I'm doing wrong, but can someone explain how to create training and testing sets off the public leaderboard data? I've been trying to use that code on Ben Hammer's GitHub repository, but it seems like I have to string a ton of little\r\n hacks together just to get it to produce some data--data that I doubt is correct probably due to one of my hacks that forced it to work.</p>\r\n<p>Thanks,</p>\r\n<p>-Brett</p>\r\n<p>[/quote]Sorry for any confusion. The public leaderboard data is a test set, and you should be making / submitting predictions based on that (not trying to create training sets from this data - you're missing all the information after the cutoff time from\r\n each day).</p>\r\n<p>When time permits, I'll clean up the Github code and release a tutorial on deriving a test set from the training data that you can use for internal model development or validation.</p>\r\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19043",
      "postDate": "01/06/2013 20:39:54",
      "content": "<p>Ben:<br>\r\nit would help that instead of writing code, if you gave some pointers on the best way to do Cross Validation for this exercise.<br>\r\nThat would be the greatest help</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "19120",
      "postDate": "01/09/2013 06:51:00",
      "content": "<p>Ben:</p>\r\n<p>No need for code. Just tell us the logic.</p>\r\n<p>It has become very difficult to create a test set from training set</p>\r\n<p>&nbsp;</p>\r\n<p>Thanks<br>\r\nkiran&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>[quote=Ben Hamner;18610]</p>\r\n<p>[quote=Brett;18556]</p>\r\n<p>Hey, I don't know what I'm doing wrong, but can someone explain how to create training and testing sets off the public leaderboard data? I've been trying to use that code on Ben Hammer's GitHub repository, but it seems like I have to string a ton of little\r\n hacks together just to get it to produce some data--data that I doubt is correct probably due to one of my hacks that forced it to work.</p>\r\n<p>Thanks,</p>\r\n<p>-Brett</p>\r\n<p>[/quote]Sorry for any confusion. The public leaderboard data is a test set, and you should be making / submitting predictions based on that (not trying to create training sets from this data - you're missing all the information after the cutoff time from\r\n each day).</p>\r\n<p>When time permits, I'll clean up the Github code and release a tutorial on deriving a test set from the training data that you can use for internal model development or validation.</p>\r\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 18610,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "12/26/2012 06:41:18",
      "content": "<p>[quote=Brett;18556]</p>\r\n<p>Hey, I don't know what I'm doing wrong, but can someone explain how to create training and testing sets off the public leaderboard data? I've been trying to use that code on Ben Hammer's GitHub repository, but it seems like I have to string a ton of little\r\n hacks together just to get it to produce some data--data that I doubt is correct probably due to one of my hacks that forced it to work.</p>\r\n<p>Thanks,</p>\r\n<p>-Brett</p>\r\n<p>[/quote]Sorry for any confusion. The public leaderboard data is a test set, and you should be making / submitting predictions based on that (not trying to create training sets from this data - you're missing all the information after the cutoff time from\r\n each day).</p>\r\n<p>When time permits, I'll clean up the Github code and release a tutorial on deriving a test set from the training data that you can use for internal model development or validation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 19032,
      "author_name": "albeam",
      "author_url": "",
      "post_date": "01/06/2013 16:54:33",
      "content": "<p>[quote=Ben Hamner;18610]</p>\r\n<p>[quote=Brett;18556]</p>\r\n<p>Hey, I don't know what I'm doing wrong, but can someone explain how to create training and testing sets off the public leaderboard data? I've been trying to use that code on Ben Hammer's GitHub repository, but it seems like I have to string a ton of little\r\n hacks together just to get it to produce some data--data that I doubt is correct probably due to one of my hacks that forced it to work.</p>\r\n<p>Thanks,</p>\r\n<p>-Brett</p>\r\n<p>[/quote]Sorry for any confusion. The public leaderboard data is a test set, and you should be making / submitting predictions based on that (not trying to create training sets from this data - you're missing all the information after the cutoff time from\r\n each day).</p>\r\n<p>When time permits, I'll clean up the Github code and release a tutorial on deriving a test set from the training data that you can use for internal model development or validation.</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>Has this tutorial ever been released?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 19035,
      "author_name": "rickyars",
      "author_url": "",
      "post_date": "01/06/2013 18:02:43",
      "content": "<p>I think there are 3-5 forum topics all asking variations of the same question. It would be nice if we could get conirmation that we will be given a list of flight_history_ids to predict for...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 19036,
      "author_name": "vadnais",
      "author_url": "",
      "post_date": "01/06/2013 18:10:39",
      "content": "<p>If you read the Submission Instructions at:</p>\r\n<p><a href=\"http://www.gequest.com/c/flight/details/submission-instructions\">http://www.gequest.com/c/flight/details/submission-instructions</a></p>\r\n<p>you will see that:</p>\r\n<p>“Sample submission files (such as estimated_arrival_benchmark.csv) for the public&nbsp;leaderboard set have been uploaded to the\r\n<a href=\"https://www.gequest.com/c/flight/data\">data page</a>. A sample submission for the final evaluation set will be released with that set.”</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 19040,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "01/06/2013 19:55:41",
      "content": "<p>ok, look forward to it</p>\r\n<p>&nbsp;</p>\r\n<p>[quote=Ben Hamner;18610]</p>\r\n<p>[quote=Brett;18556]</p>\r\n<p>Hey, I don't know what I'm doing wrong, but can someone explain how to create training and testing sets off the public leaderboard data? I've been trying to use that code on Ben Hammer's GitHub repository, but it seems like I have to string a ton of little\r\n hacks together just to get it to produce some data--data that I doubt is correct probably due to one of my hacks that forced it to work.</p>\r\n<p>Thanks,</p>\r\n<p>-Brett</p>\r\n<p>[/quote]Sorry for any confusion. The public leaderboard data is a test set, and you should be making / submitting predictions based on that (not trying to create training sets from this data - you're missing all the information after the cutoff time from\r\n each day).</p>\r\n<p>When time permits, I'll clean up the Github code and release a tutorial on deriving a test set from the training data that you can use for internal model development or validation.</p>\r\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 19043,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "01/06/2013 20:39:54",
      "content": "<p>Ben:<br>\r\nit would help that instead of writing code, if you gave some pointers on the best way to do Cross Validation for this exercise.<br>\r\nThat would be the greatest help</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 19120,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "01/09/2013 06:51:00",
      "content": "<p>Ben:</p>\r\n<p>No need for code. Just tell us the logic.</p>\r\n<p>It has become very difficult to create a test set from training set</p>\r\n<p>&nbsp;</p>\r\n<p>Thanks<br>\r\nkiran&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>[quote=Ben Hamner;18610]</p>\r\n<p>[quote=Brett;18556]</p>\r\n<p>Hey, I don't know what I'm doing wrong, but can someone explain how to create training and testing sets off the public leaderboard data? I've been trying to use that code on Ben Hammer's GitHub repository, but it seems like I have to string a ton of little\r\n hacks together just to get it to produce some data--data that I doubt is correct probably due to one of my hacks that forced it to work.</p>\r\n<p>Thanks,</p>\r\n<p>-Brett</p>\r\n<p>[/quote]Sorry for any confusion. The public leaderboard data is a test set, and you should be making / submitting predictions based on that (not trying to create training sets from this data - you're missing all the information after the cutoff time from\r\n each day).</p>\r\n<p>When time permits, I'll clean up the Github code and release a tutorial on deriving a test set from the training data that you can use for internal model development or validation.</p>\r\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "18556": "",
    "18610": "",
    "19032": "",
    "19035": "",
    "19036": "",
    "19040": "",
    "19043": "",
    "19120": ""
  },
  "source": "meta"
}