{
  "id": 20488,
  "title": "beating the benchmark in scala on a spark cluster",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20488",
  "author_name": "vtKMH",
  "post_date": "2016-04-27T18:08:42.050000",
  "votes": 3,
  "comment_count": 0,
  "views": 608,
  "content": "<p>I punted the data up to S3 last night and just tried to mimic the first beating the benchmark approach in scala on a spark cluster on aws.  This script runs in under a minute in spark-shell on a three 2xl machine cluster.</p>\n\n<p>I haven't bothered doing much with it yet, and haven't broken up the test data yet, but since it's a high bias prediction, i ran it against the training set (yes, I know) and it came out around 0.29, which seems about expected.</p>\n\n<p>Again, just a first quick and dirty pass to start examining the data, but thought it might be interesting for someone new to spark/scala.</p>\n\n<p>kevin</p>",
  "messages": [
    {
      "id": 117180,
      "postDate": "2016-04-27T18:08:42.050Z",
      "content": "<p>I punted the data up to S3 last night and just tried to mimic the first beating the benchmark approach in scala on a spark cluster on aws.  This script runs in under a minute in spark-shell on a three 2xl machine cluster.</p>\n\n<p>I haven't bothered doing much with it yet, and haven't broken up the test data yet, but since it's a high bias prediction, i ran it against the training set (yes, I know) and it came out around 0.29, which seems about expected.</p>\n\n<p>Again, just a first quick and dirty pass to start examining the data, but thought it might be interesting for someone new to spark/scala.</p>\n\n<p>kevin</p>",
      "rawMarkdown": "I punted the data up to S3 last night and just tried to mimic the first beating the benchmark approach in scala on a spark cluster on aws.  This script runs in under a minute in spark-shell on a three 2xl machine cluster.\r\n\r\nI haven't bothered doing much with it yet, and haven't broken up the test data yet, but since it's a high bias prediction, i ran it against the training set (yes, I know) and it came out around 0.29, which seems about expected.\r\n\r\nAgain, just a first quick and dirty pass to start examining the data, but thought it might be interesting for someone new to spark/scala.\r\n\r\nkevin",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "117180": "I punted the data up to S3 last night and just tried to mimic the first beating the benchmark approach in scala on a spark cluster on aws.  This script runs in under a minute in spark-shell on a three 2xl machine cluster.\r\n\r\nI haven't bothered doing much with it yet, and haven't broken up the test data yet, but since it's a high bias prediction, i ran it against the training set (yes, I know) and it came out around 0.29, which seems about expected.\r\n\r\nAgain, just a first quick and dirty pass to start examining the data, but thought it might be interesting for someone new to spark/scala.\r\n\r\nkevin"
  }
}