{
  "id": 24489,
  "title": "How To? : Utilize local compute power on the large datasets",
  "url": "/competitions/outbrain-click-prediction/discussion/24489",
  "author_name": "",
  "post_date": "2016-10-17T02:42:16.853Z",
  "votes": null,
  "comment_count": 4,
  "views": 594,
  "content": "<p>I am getting started on data science and currently focusing on performing EDAs. Getting really stuck up with large datasets and am using a macbook with 8GB ram; Currently limited to local compute horse power. </p>\n\n<p>Any thoughts/references on how to handle them would be very helpful. Outbrains dataset looks a ideal fit to get this fixed :)</p>\n\n<p><strong>Example:</strong> </p>\n\n<ul>\n<li>A Merge using R (in RStudio) on <em>Clicks</em> and <em>Promoted Content</em> takes hell lot of time though the compute power is fully dedicated for this process.</li>\n<li>A simple grouby on the above merged dataset makes a graceful kernel death with out of memory in jupyter notebook using python.</li>\n</ul>\n\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": "139874",
      "postDate": "10/17/2016 02:42:16",
      "content": "<p>I am getting started on data science and currently focusing on performing EDAs. Getting really stuck up with large datasets and am using a macbook with 8GB ram; Currently limited to local compute horse power. </p>\n\n<p>Any thoughts/references on how to handle them would be very helpful. Outbrains dataset looks a ideal fit to get this fixed :)</p>\n\n<p><strong>Example:</strong> </p>\n\n<ul>\n<li>A Merge using R (in RStudio) on <em>Clicks</em> and <em>Promoted Content</em> takes hell lot of time though the compute power is fully dedicated for this process.</li>\n<li>A simple grouby on the above merged dataset makes a graceful kernel death with out of memory in jupyter notebook using python.</li>\n</ul>\n\n<p>Thanks in advance!</p>",
      "rawMarkdown": "I am getting started on data science and currently focusing on performing EDAs. Getting really stuck up with large datasets and am using a macbook with 8GB ram; Currently limited to local compute horse power. \r\n\r\nAny thoughts/references on how to handle them would be very helpful. Outbrains dataset looks a ideal fit to get this fixed :)\r\n\r\n**Example:** \r\n\r\n- A Merge using R (in RStudio) on *Clicks* and *Promoted Content* takes hell lot of time though the compute power is fully dedicated for this process.\r\n- A simple grouby on the above merged dataset makes a graceful kernel death with out of memory in jupyter notebook using python.\r\n\r\nThanks in advance!",
      "votes": null
    },
    {
      "id": "140017",
      "postDate": "10/17/2016 23:04:19",
      "content": "<p>One approach may be using a database for joins and group bys.\nNot all of them are good at out-of-memory operations, though.</p>\n\n<p>Or you can try creating your own, something like <a href=\"http://www.vldb.org/pvldb/vol8/p353-barber.pdf\">Memory-Efficient Hash Joins</a>, or many more others. Doing joins and group-bys for data sets larger than memory used to be a hot research topic when big RAM \nmeant 8MB :)</p>",
      "rawMarkdown": "One approach may be using a database for joins and group bys.\r\nNot all of them are good at out-of-memory operations, though.\r\n\r\nOr you can try creating your own, something like [Memory-Efficient Hash Joins][1], or many more others. Doing joins and group-bys for data sets larger than memory used to be a hot research topic when big RAM \r\nmeant 8MB :)\r\n\r\n  [1]: http://www.vldb.org/pvldb/vol8/p353-barber.pdf",
      "votes": null
    },
    {
      "id": "140026",
      "postDate": "10/18/2016 00:59:03",
      "content": "<p>You can use distributed processing for joining and merging large datasets that do not fit in memory. The more popular distributed processing platform is Hadoop. You can easily start a Hadoop cluster on <a href=\"https://aws.amazon.com/emr/?nc2=h_m1\">Amazon EMR</a>, <a href=\"http://cloud.google.com/dataproc/\">Google Dataproc</a> or <a href=\"https://azure.microsoft.com/en-us/documentation/articles/hdinsight-apache-spark-overview/\">Azure HDInsight</a>. <br>\nI suggest you to use on top of Hadoop a higher level querying interface, like Apache Hive (SQL) or Spark SQL (Dataframes like Python-Pandas or R).\nIf you want to use Hive, for example, first you need to just upload your CSV datasets to Hadoop HDFS with commands like below:</p>\n\n<pre><code>hadoop fs -mkdir outbrain/page_views  \nhadoop fs -put ./page_views.csv outbrain/page_views\n</code></pre>\n\n<p>Then you can map on Hive an external table based on the CSV file, for example:</p>\n\n<pre><code>CREATE EXTERNAL TABLE IF NOT EXISTS PAGE_VIEWS(\n        uuid STRING, \n        document_id INT,\n        time_stamp INT,\n        platform INT,\n        geo_location STRING, \n        traffic_source INT\n    COMMENT 'Data from page_views.csv')\n    ROW FORMAT DELIMITED\n    FIELDS TERMINATED BY ','\n    STORED AS TEXTFILE\n    location '/user/root/outbrain/page_views' -- Folder on HDFS where is your CSV file\n    tblproperties (&quot;skip.header.line.count&quot;=&quot;1&quot;);\n</code></pre>\n\n<p>Now you can map your other CSV tables on Hive and use vanilla SQL (or HiveQL) to perform you queries, and joins.</p>\n\n<p>If you prefer the DataFrame approach, take a look on this <a href=\"https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/\">kernel where I explore page_views.csv using PySpark</a>.</p>",
      "rawMarkdown": "You can use distributed processing for joining and merging large datasets that do not fit in memory. The more popular distributed processing platform is Hadoop. You can easily start a Hadoop cluster on [Amazon EMR](https://aws.amazon.com/emr/?nc2=h_m1), [Google Dataproc](http://cloud.google.com/dataproc/) or [Azure HDInsight](https://azure.microsoft.com/en-us/documentation/articles/hdinsight-apache-spark-overview/).  \r\nI suggest you to use on top of Hadoop a higher level querying interface, like Apache Hive (SQL) or Spark SQL (Dataframes like Python-Pandas or R).\r\nIf you want to use Hive, for example, first you need to just upload your CSV datasets to Hadoop HDFS with commands like below:\r\n\r\n    hadoop fs -mkdir outbrain/page_views  \r\n    hadoop fs -put ./page_views.csv outbrain/page_views\r\n\r\nThen you can map on Hive an external table based on the CSV file, for example:\r\n\r\n    CREATE EXTERNAL TABLE IF NOT EXISTS PAGE_VIEWS(\r\n            uuid STRING, \r\n            document_id INT,\r\n            time_stamp INT,\r\n            platform INT,\r\n            geo_location STRING, \r\n            traffic_source INT\r\n        COMMENT 'Data from page_views.csv')\r\n        ROW FORMAT DELIMITED\r\n        FIELDS TERMINATED BY ','\r\n        STORED AS TEXTFILE\r\n        location '/user/root/outbrain/page_views' -- Folder on HDFS where is your CSV file\r\n        tblproperties (\"skip.header.line.count\"=\"1\");\r\n\r\nNow you can map your other CSV tables on Hive and use vanilla SQL (or HiveQL) to perform you queries, and joins.\r\n\r\nIf you prefer the DataFrame approach, take a look on this [kernel where I explore page_views.csv using PySpark](https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/).",
      "votes": null
    },
    {
      "id": "140044",
      "postDate": "10/18/2016 02:24:31",
      "content": "<blockquote>\n  <p>You can use distributed processing</p>\n</blockquote>\n\n<p>I thought the OP specifically asked for <em>local</em> solution, which a distributed one is not :)</p>",
      "rawMarkdown": "> You can use distributed processing\r\n\r\nI thought the OP specifically asked for _local_ solution, which a distributed one is not :)",
      "votes": null
    },
    {
      "id": "140213",
      "postDate": "10/18/2016 23:07:20",
      "content": "<p>You could try using <a href=\"https://turi.com/\">GraphLab</a>, which has an out-of-core  implementation of data management and algorithms.</p>\n\n<p>While not a local soln, I find using AWS and just using a m2.xlarge  (8 core/30GB) or larger machine is the simplest solution. Cost is about 8c/hr</p>",
      "rawMarkdown": "You could try using [GraphLab][1], which has an out-of-core  implementation of data management and algorithms.\r\n\r\nWhile not a local soln, I find using AWS and just using a m2.xlarge  (8 core/30GB) or larger machine is the simplest solution. Cost is about 8c/hr\r\n\r\n\r\n  [1]: https://turi.com/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 140017,
      "author_name": "andris",
      "author_url": "",
      "post_date": "10/17/2016 23:04:19",
      "content": "<p>One approach may be using a database for joins and group bys.\nNot all of them are good at out-of-memory operations, though.</p>\n\n<p>Or you can try creating your own, something like <a href=\"http://www.vldb.org/pvldb/vol8/p353-barber.pdf\">Memory-Efficient Hash Joins</a>, or many more others. Doing joins and group-bys for data sets larger than memory used to be a hot research topic when big RAM \nmeant 8MB :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 140026,
      "author_name": "gspmoreira",
      "author_url": "",
      "post_date": "10/18/2016 00:59:03",
      "content": "<p>You can use distributed processing for joining and merging large datasets that do not fit in memory. The more popular distributed processing platform is Hadoop. You can easily start a Hadoop cluster on <a href=\"https://aws.amazon.com/emr/?nc2=h_m1\">Amazon EMR</a>, <a href=\"http://cloud.google.com/dataproc/\">Google Dataproc</a> or <a href=\"https://azure.microsoft.com/en-us/documentation/articles/hdinsight-apache-spark-overview/\">Azure HDInsight</a>. <br>\nI suggest you to use on top of Hadoop a higher level querying interface, like Apache Hive (SQL) or Spark SQL (Dataframes like Python-Pandas or R).\nIf you want to use Hive, for example, first you need to just upload your CSV datasets to Hadoop HDFS with commands like below:</p>\n\n<pre><code>hadoop fs -mkdir outbrain/page_views  \nhadoop fs -put ./page_views.csv outbrain/page_views\n</code></pre>\n\n<p>Then you can map on Hive an external table based on the CSV file, for example:</p>\n\n<pre><code>CREATE EXTERNAL TABLE IF NOT EXISTS PAGE_VIEWS(\n        uuid STRING, \n        document_id INT,\n        time_stamp INT,\n        platform INT,\n        geo_location STRING, \n        traffic_source INT\n    COMMENT 'Data from page_views.csv')\n    ROW FORMAT DELIMITED\n    FIELDS TERMINATED BY ','\n    STORED AS TEXTFILE\n    location '/user/root/outbrain/page_views' -- Folder on HDFS where is your CSV file\n    tblproperties (&quot;skip.header.line.count&quot;=&quot;1&quot;);\n</code></pre>\n\n<p>Now you can map your other CSV tables on Hive and use vanilla SQL (or HiveQL) to perform you queries, and joins.</p>\n\n<p>If you prefer the DataFrame approach, take a look on this <a href=\"https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/\">kernel where I explore page_views.csv using PySpark</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 140044,
      "author_name": "andris",
      "author_url": "",
      "post_date": "10/18/2016 02:24:31",
      "content": "<blockquote>\n  <p>You can use distributed processing</p>\n</blockquote>\n\n<p>I thought the OP specifically asked for <em>local</em> solution, which a distributed one is not :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 140213,
      "author_name": "kevinmcisaac",
      "author_url": "",
      "post_date": "10/18/2016 23:07:20",
      "content": "<p>You could try using <a href=\"https://turi.com/\">GraphLab</a>, which has an out-of-core  implementation of data management and algorithms.</p>\n\n<p>While not a local soln, I find using AWS and just using a m2.xlarge  (8 core/30GB) or larger machine is the simplest solution. Cost is about 8c/hr</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "139874": "I am getting started on data science and currently focusing on performing EDAs. Getting really stuck up with large datasets and am using a macbook with 8GB ram; Currently limited to local compute horse power. \r\n\r\nAny thoughts/references on how to handle them would be very helpful. Outbrains dataset looks a ideal fit to get this fixed :)\r\n\r\n**Example:** \r\n\r\n- A Merge using R (in RStudio) on *Clicks* and *Promoted Content* takes hell lot of time though the compute power is fully dedicated for this process.\r\n- A simple grouby on the above merged dataset makes a graceful kernel death with out of memory in jupyter notebook using python.\r\n\r\nThanks in advance!",
    "140017": "One approach may be using a database for joins and group bys.\r\nNot all of them are good at out-of-memory operations, though.\r\n\r\nOr you can try creating your own, something like [Memory-Efficient Hash Joins][1], or many more others. Doing joins and group-bys for data sets larger than memory used to be a hot research topic when big RAM \r\nmeant 8MB :)\r\n\r\n  [1]: http://www.vldb.org/pvldb/vol8/p353-barber.pdf",
    "140026": "You can use distributed processing for joining and merging large datasets that do not fit in memory. The more popular distributed processing platform is Hadoop. You can easily start a Hadoop cluster on [Amazon EMR](https://aws.amazon.com/emr/?nc2=h_m1), [Google Dataproc](http://cloud.google.com/dataproc/) or [Azure HDInsight](https://azure.microsoft.com/en-us/documentation/articles/hdinsight-apache-spark-overview/).  \r\nI suggest you to use on top of Hadoop a higher level querying interface, like Apache Hive (SQL) or Spark SQL (Dataframes like Python-Pandas or R).\r\nIf you want to use Hive, for example, first you need to just upload your CSV datasets to Hadoop HDFS with commands like below:\r\n\r\n    hadoop fs -mkdir outbrain/page_views  \r\n    hadoop fs -put ./page_views.csv outbrain/page_views\r\n\r\nThen you can map on Hive an external table based on the CSV file, for example:\r\n\r\n    CREATE EXTERNAL TABLE IF NOT EXISTS PAGE_VIEWS(\r\n            uuid STRING, \r\n            document_id INT,\r\n            time_stamp INT,\r\n            platform INT,\r\n            geo_location STRING, \r\n            traffic_source INT\r\n        COMMENT 'Data from page_views.csv')\r\n        ROW FORMAT DELIMITED\r\n        FIELDS TERMINATED BY ','\r\n        STORED AS TEXTFILE\r\n        location '/user/root/outbrain/page_views' -- Folder on HDFS where is your CSV file\r\n        tblproperties (\"skip.header.line.count\"=\"1\");\r\n\r\nNow you can map your other CSV tables on Hive and use vanilla SQL (or HiveQL) to perform you queries, and joins.\r\n\r\nIf you prefer the DataFrame approach, take a look on this [kernel where I explore page_views.csv using PySpark](https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/).",
    "140044": "> You can use distributed processing\r\n\r\nI thought the OP specifically asked for _local_ solution, which a distributed one is not :)",
    "140213": "You could try using [GraphLab][1], which has an out-of-core  implementation of data management and algorithms.\r\n\r\nWhile not a local soln, I find using AWS and just using a m2.xlarge  (8 core/30GB) or larger machine is the simplest solution. Cost is about 8c/hr\r\n\r\n\r\n  [1]: https://turi.com/"
  },
  "source": "meta"
}