{
  "id": 3953,
  "title": "Where to start?",
  "url": "/competitions/flight/discussion/3953",
  "author_name": "",
  "post_date": "2013-03-03T07:03:15.237Z",
  "votes": null,
  "comment_count": 4,
  "views": 7102,
  "content": "<p>Hey,</p>\r\n<p>I am interested in doing this for educational purposes and I doubt I will be a &quot;winner&quot; in this competition. &nbsp;I would like to know how one should begin when dealing with datasets so large. &nbsp;Do you do some preliminary work with data-plotting or go straight\r\n to the random forests algorithms? &nbsp;I would like to hear your approach. &nbsp;</p>",
  "messages": [
    {
      "id": "21013",
      "postDate": "03/03/2013 07:03:15",
      "content": "<p>Hey,</p>\r\n<p>I am interested in doing this for educational purposes and I doubt I will be a &quot;winner&quot; in this competition. &nbsp;I would like to know how one should begin when dealing with datasets so large. &nbsp;Do you do some preliminary work with data-plotting or go straight\r\n to the random forests algorithms? &nbsp;I would like to hear your approach. &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "21023",
      "postDate": "03/03/2013 19:28:48",
      "content": "<p>Hey Dan,</p>\r\n<p><span style=\"line-height:1.4em\">I'm just starting out too, and thought this would be an interesting one ot start with, I had some ideas about where to begin. I'm going to extract the data and basically look at afew instances of flights history and brgin\r\n making some small basic groupings, so identifying schdeuled flights, looking at certain routes, and get a really good feel for the data first with some basic visualisations.\r\n</span></p>\r\n<p><span style=\"line-height:1.4em\">If you are interested we could team up and share insights/models. I am going to use Qlikview for the basic visualisation stuff, mapping etc and R/Python for the modeling. I was going to rent an EC2 for the crunching. I'm based\r\n in the UK and will be pottering on this in the evenings.</span></p>\r\n<p><span style=\"line-height:1.4em\">Let me know</span></p>\r\n<p><span style=\"line-height:1.4em\">Mark&nbsp;</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "21024",
      "postDate": "03/03/2013 20:17:25",
      "content": "<p>Mark,</p>\r\n<p>I would be interested but I don't know how much I can commit to it. &nbsp;Email me with some ideas and we can get started.</p>\r\n<p>&nbsp;</p>\r\n<p>Dan</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "21025",
      "postDate": "03/03/2013 21:16:06",
      "content": "<p>In our case 95% of the work was done in SQL. The data itself is very &quot;buggy&quot; and whole system must be very fault tolerant. Thus when we had an overview of the data we looked into reducing the biggest errors for example those above 60 minutes. The error metric\r\n is very prone to such errors but they are rather easy to catch.&nbsp;</p>\r\n<p>Our weak model was about ~6.5 on the leaderboard. So start with simple models that you can treat as a benchmark. If you cannot achieve such result with a simple model there is no sense going further (in terms of model complexity). To sum up expect spending\r\n more than 90% of the time on the data extraction and cleaning.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "39492",
      "postDate": "02/26/2014 05:50:58",
      "content": "<p>[quote=Daniel Parry;21013]</p>\n<p>Hey,</p>\n<p>I am interested in doing this for educational purposes and I doubt I will be a &quot;winner&quot; in this competition. &nbsp;I would like to know how one should begin when dealing with datasets so large. &nbsp;Do you do some preliminary work with data-plotting or go straight to the random forests algorithms? &nbsp;I would like to hear your approach. &nbsp;</p>\n<p>[/quote]</p>\n<p>Hey&#65292;Dan</p>\n<p>I am very&nbsp;interested in using&nbsp;random forests algorithms on this problem, would you give me some suggest?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 21023,
      "author_name": "macclad",
      "author_url": "",
      "post_date": "03/03/2013 19:28:48",
      "content": "<p>Hey Dan,</p>\r\n<p><span style=\"line-height:1.4em\">I'm just starting out too, and thought this would be an interesting one ot start with, I had some ideas about where to begin. I'm going to extract the data and basically look at afew instances of flights history and brgin\r\n making some small basic groupings, so identifying schdeuled flights, looking at certain routes, and get a really good feel for the data first with some basic visualisations.\r\n</span></p>\r\n<p><span style=\"line-height:1.4em\">If you are interested we could team up and share insights/models. I am going to use Qlikview for the basic visualisation stuff, mapping etc and R/Python for the modeling. I was going to rent an EC2 for the crunching. I'm based\r\n in the UK and will be pottering on this in the evenings.</span></p>\r\n<p><span style=\"line-height:1.4em\">Let me know</span></p>\r\n<p><span style=\"line-height:1.4em\">Mark&nbsp;</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 21024,
      "author_name": "dantparry",
      "author_url": "",
      "post_date": "03/03/2013 20:17:25",
      "content": "<p>Mark,</p>\r\n<p>I would be interested but I don't know how much I can commit to it. &nbsp;Email me with some ideas and we can get started.</p>\r\n<p>&nbsp;</p>\r\n<p>Dan</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 21025,
      "author_name": "paweljankiewicz",
      "author_url": "",
      "post_date": "03/03/2013 21:16:06",
      "content": "<p>In our case 95% of the work was done in SQL. The data itself is very &quot;buggy&quot; and whole system must be very fault tolerant. Thus when we had an overview of the data we looked into reducing the biggest errors for example those above 60 minutes. The error metric\r\n is very prone to such errors but they are rather easy to catch.&nbsp;</p>\r\n<p>Our weak model was about ~6.5 on the leaderboard. So start with simple models that you can treat as a benchmark. If you cannot achieve such result with a simple model there is no sense going further (in terms of model complexity). To sum up expect spending\r\n more than 90% of the time on the data extraction and cleaning.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 39492,
      "author_name": "ashqal",
      "author_url": "",
      "post_date": "02/26/2014 05:50:58",
      "content": "<p>[quote=Daniel Parry;21013]</p>\n<p>Hey,</p>\n<p>I am interested in doing this for educational purposes and I doubt I will be a &quot;winner&quot; in this competition. &nbsp;I would like to know how one should begin when dealing with datasets so large. &nbsp;Do you do some preliminary work with data-plotting or go straight to the random forests algorithms? &nbsp;I would like to hear your approach. &nbsp;</p>\n<p>[/quote]</p>\n<p>Hey&#65292;Dan</p>\n<p>I am very&nbsp;interested in using&nbsp;random forests algorithms on this problem, would you give me some suggest?&nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "21013": "",
    "21023": "",
    "21024": "",
    "21025": "",
    "39492": ""
  },
  "source": "meta"
}