{
  "id": 20896,
  "title": "Leakage solution with Spark SQL / pyspark",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20896",
  "author_name": "",
  "post_date": "2016-05-12T10:37:16.313Z",
  "votes": 18,
  "comment_count": 2,
  "views": 1886,
  "content": "<p>Hi all,</p>\n\n<p>I know many of you are struggling with the sheer volume of 4 GB, as we all know, if it crashes your Excel, it is Big Data :)</p>\n\n<p>I want to share with you all a starter script that imitates ZFTurbo's leakage solution and uses Spark to crunch the dataset and build aggregates. </p>\n\n<p><a href=\"http://nbviewer.jupyter.org/gist/vykhand/1f2484ff14fbbf805234160cf90668b4\">NBViewer link</a></p>\n\n<p><a href=\"https://gist.github.com/vykhand/1f2484ff14fbbf805234160cf90668b4\">Gist link</a></p>\n\n<p>I don't think there is a way to make it a Kaggle script, but if there is, please let me know</p>",
  "messages": [
    {
      "id": "119700",
      "postDate": "05/12/2016 10:37:16",
      "content": "<p>Hi all,</p>\n\n<p>I know many of you are struggling with the sheer volume of 4 GB, as we all know, if it crashes your Excel, it is Big Data :)</p>\n\n<p>I want to share with you all a starter script that imitates ZFTurbo's leakage solution and uses Spark to crunch the dataset and build aggregates. </p>\n\n<p><a href=\"http://nbviewer.jupyter.org/gist/vykhand/1f2484ff14fbbf805234160cf90668b4\">NBViewer link</a></p>\n\n<p><a href=\"https://gist.github.com/vykhand/1f2484ff14fbbf805234160cf90668b4\">Gist link</a></p>\n\n<p>I don't think there is a way to make it a Kaggle script, but if there is, please let me know</p>",
      "rawMarkdown": "Hi all,\r\n\r\nI know many of you are struggling with the sheer volume of 4 GB, as we all know, if it crashes your Excel, it is Big Data :)\r\n\r\nI want to share with you all a starter script that imitates ZFTurbo's leakage solution and uses Spark to crunch the dataset and build aggregates. \r\n\r\n[NBViewer link][1]\r\n\r\n\r\n[Gist link][2]\r\n\r\nI don't think there is a way to make it a Kaggle script, but if there is, please let me know\r\n\r\n  [1]: http://nbviewer.jupyter.org/gist/vykhand/1f2484ff14fbbf805234160cf90668b4\r\n  [2]: https://gist.github.com/vykhand/1f2484ff14fbbf805234160cf90668b4",
      "votes": null
    },
    {
      "id": "122256",
      "postDate": "06/02/2016 14:20:41",
      "content": "<p>Hi,</p>\n\n<p>Meanwhile I've ported your code to Scala Spark here.\n<a href=\"https://github.com/cupuyc/scala-kaggle/blob/master/src/main/scala/io/github/stanreshetnyk/expedia/ExpediaSpark.scala\">https://github.com/cupuyc/scala-kaggle/blob/master/src/main/scala/io/github/stanreshetnyk/expedia/ExpediaSpark.scala</a></p>\n\n<p>It was useful for me to get more experience with Spark and Scala.</p>\n\n<p>Stan</p>",
      "rawMarkdown": "Hi,\r\n\r\nMeanwhile I've ported your code to Scala Spark here.\r\nhttps://github.com/cupuyc/scala-kaggle/blob/master/src/main/scala/io/github/stanreshetnyk/expedia/ExpediaSpark.scala\r\n\r\nIt was useful for me to get more experience with Spark and Scala.\r\n\r\nStan",
      "votes": null
    },
    {
      "id": "122301",
      "postDate": "06/02/2016 21:29:27",
      "content": "<p>@Stan, excellent, thanks for sharing! I will use it as a reference comparing Spark Scala code and pyspark code. </p>",
      "rawMarkdown": "Stan, excellent, thanks for sharing! I will use it as a reference comparing Spark Scala code and pyspark code.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 122256,
      "author_name": "cupuyc",
      "author_url": "",
      "post_date": "06/02/2016 14:20:41",
      "content": "<p>Hi,</p>\n\n<p>Meanwhile I've ported your code to Scala Spark here.\n<a href=\"https://github.com/cupuyc/scala-kaggle/blob/master/src/main/scala/io/github/stanreshetnyk/expedia/ExpediaSpark.scala\">https://github.com/cupuyc/scala-kaggle/blob/master/src/main/scala/io/github/stanreshetnyk/expedia/ExpediaSpark.scala</a></p>\n\n<p>It was useful for me to get more experience with Spark and Scala.</p>\n\n<p>Stan</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122301,
      "author_name": "vykhand",
      "author_url": "",
      "post_date": "06/02/2016 21:29:27",
      "content": "<p>@Stan, excellent, thanks for sharing! I will use it as a reference comparing Spark Scala code and pyspark code. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119700": "Hi all,\r\n\r\nI know many of you are struggling with the sheer volume of 4 GB, as we all know, if it crashes your Excel, it is Big Data :)\r\n\r\nI want to share with you all a starter script that imitates ZFTurbo's leakage solution and uses Spark to crunch the dataset and build aggregates. \r\n\r\n[NBViewer link][1]\r\n\r\n\r\n[Gist link][2]\r\n\r\nI don't think there is a way to make it a Kaggle script, but if there is, please let me know\r\n\r\n  [1]: http://nbviewer.jupyter.org/gist/vykhand/1f2484ff14fbbf805234160cf90668b4\r\n  [2]: https://gist.github.com/vykhand/1f2484ff14fbbf805234160cf90668b4",
    "122256": "Hi,\r\n\r\nMeanwhile I've ported your code to Scala Spark here.\r\nhttps://github.com/cupuyc/scala-kaggle/blob/master/src/main/scala/io/github/stanreshetnyk/expedia/ExpediaSpark.scala\r\n\r\nIt was useful for me to get more experience with Spark and Scala.\r\n\r\nStan",
    "122301": "Stan, excellent, thanks for sharing! I will use it as a reference comparing Spark Scala code and pyspark code."
  },
  "source": "meta"
}