{
  "id": 39858,
  "title": "Try using Spark !",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/39858",
  "author_name": "",
  "post_date": "2017-09-22T13:05:06.137389900Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I used Spark to perform EDA and preprocessing on the entire data.</p>\n\n<p>It takes less than 30 seconds per graph, Try using Spark !</p>\n\n<p>Is there a way to put zeppelin notebook in the kernel?</p>",
  "messages": [
    {
      "id": "223514",
      "postDate": "09/22/2017 13:05:06",
      "content": "<p>I used Spark to perform EDA and preprocessing on the entire data.</p>\n\n<p>It takes less than 30 seconds per graph, Try using Spark !</p>\n\n<p>Is there a way to put zeppelin notebook in the kernel?</p>",
      "rawMarkdown": "I used Spark to perform EDA and preprocessing on the entire data.\n\nIt takes less than 30 seconds per graph, Try using Spark !\n\nIs there a way to put zeppelin notebook in the kernel?",
      "votes": null
    },
    {
      "id": "223540",
      "postDate": "09/22/2017 14:13:28",
      "content": "<p>Great!  Does it work  in local mode?    Whats your hardware environment? Thanks</p>",
      "rawMarkdown": "Great!  Does it work  in local mode?    Whats your hardware environment? Thanks",
      "votes": null
    },
    {
      "id": "223678",
      "postDate": "09/23/2017 02:39:36",
      "content": "<p>Sure, but Spark is a distributed processing framework.\nAdding worker nodes makes execution time faster.\nI use my AWS EMR cluster.</p>\n\n<ul>\n<li>Master : m4.xlarge, num = 1</li>\n<li>Worker : m4.xlarge, num = 3 (spot instances)</li>\n</ul>",
      "rawMarkdown": "Sure, but Spark is a distributed processing framework.\nAdding worker nodes makes execution time faster.\nI use my AWS EMR cluster.\n\n - Master : m4.xlarge, num = 1\n - Worker : m4.xlarge, num = 3 (spot instances)",
      "votes": null
    },
    {
      "id": "224307",
      "postDate": "09/25/2017 21:34:21",
      "content": "<p>@JunyoungPark - how are you managing the cluster and provisioning the instances etc.? Thanks :)</p>",
      "rawMarkdown": "JunyoungPark - how are you managing the cluster and provisioning the instances etc.? Thanks :)",
      "votes": null
    },
    {
      "id": "224440",
      "postDate": "09/26/2017 10:35:47",
      "content": "<p>@Daniel Burkhardt Cerigo - \nAWS EMR addresses everything that need to manage cluster. EMR can build clusters with just click of a button, so we can focus more on competitions. If EMR is expensive, Cloudera Manager on EC2 might be a good choice.</p>",
      "rawMarkdown": "Daniel Burkhardt Cerigo - \nAWS EMR addresses everything that need to manage cluster. EMR can build clusters with just click of a button, so we can focus more on competitions. If EMR is expensive, Cloudera Manager on EC2 might be a good choice.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 223540,
      "author_name": "",
      "author_url": "",
      "post_date": "09/22/2017 14:13:28",
      "content": "<p>Great!  Does it work  in local mode?    Whats your hardware environment? Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 223678,
          "author_name": "junyoung",
          "author_url": "",
          "post_date": "09/23/2017 02:39:36",
          "content": "<p>Sure, but Spark is a distributed processing framework.\nAdding worker nodes makes execution time faster.\nI use my AWS EMR cluster.</p>\n\n<ul>\n<li>Master : m4.xlarge, num = 1</li>\n<li>Worker : m4.xlarge, num = 3 (spot instances)</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 224307,
      "author_name": "dbcerigo",
      "author_url": "",
      "post_date": "09/25/2017 21:34:21",
      "content": "<p>@JunyoungPark - how are you managing the cluster and provisioning the instances etc.? Thanks :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 224440,
          "author_name": "junyoung",
          "author_url": "",
          "post_date": "09/26/2017 10:35:47",
          "content": "<p>@Daniel Burkhardt Cerigo - \nAWS EMR addresses everything that need to manage cluster. EMR can build clusters with just click of a button, so we can focus more on competitions. If EMR is expensive, Cloudera Manager on EC2 might be a good choice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "223514": "I used Spark to perform EDA and preprocessing on the entire data.\n\nIt takes less than 30 seconds per graph, Try using Spark !\n\nIs there a way to put zeppelin notebook in the kernel?",
    "223540": "Great!  Does it work  in local mode?    Whats your hardware environment? Thanks",
    "223678": "Sure, but Spark is a distributed processing framework.\nAdding worker nodes makes execution time faster.\nI use my AWS EMR cluster.\n\n - Master : m4.xlarge, num = 1\n - Worker : m4.xlarge, num = 3 (spot instances)",
    "224307": "JunyoungPark - how are you managing the cluster and provisioning the instances etc.? Thanks :)",
    "224440": "Daniel Burkhardt Cerigo - \nAWS EMR addresses everything that need to manage cluster. EMR can build clusters with just click of a button, so we can focus more on competitions. If EMR is expensive, Cloudera Manager on EC2 might be a good choice."
  },
  "source": "meta"
}