{
  "id": 54030,
  "title": "Analysing Full Dataset with Apache Spark (local mode or not)",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54030",
  "author_name": "Meta Learner",
  "post_date": "2018-04-08T16:11:19.533000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I am opening this topic for those who are interested in leveraging the full potential of TalkingData dataset with big data tools such as Apache Spark. </p>\n\n<p>Please could you share your tips?</p>\n\n<p>It can be for example the way you are partitioning the data to circumvent memory limitations, installation processes or any other tricks that can be useful for \"Spark Beginners\" to explore Talking Data full set and perhaps help them in doing some predictions. </p>\n\n<p>I have installed Pyspark 2.1.0 on a windows 10 computer (16 GB RAM and 4 cores) with anaconda Python 3.5 and Jupyter. Here I am listing the best resources that I have found so far:</p>\n\n<ol>\n<li>Installing Pyspark on windows: <a href=\"https://medium.com/@GalarnykMichael/install-spark-on-windows-pyspark-4498a5d8d66c\">https://medium.com/@GalarnykMichael/install-spark-on-windows-pyspark-4498a5d8d66c</a></li>\n<li>Using dataframes in Pyspark: <a href=\"https://www.analyticsvidhya.com/blog/2016/10/spark-dataframe-and-operations/\">https://www.analyticsvidhya.com/blog/2016/10/spark-dataframe-and-operations/</a></li>\n<li>A classification problem handle in Spark with summary statistics: <a href=\"https://mapr.com/blog/churn-prediction-pyspark-using-mllib-and-ml-packages/\">https://mapr.com/blog/churn-prediction-pyspark-using-mllib-and-ml-packages/</a></li>\n<li>Another machine learning example in Spark +statistics+ how to apply SQL queries on spark dataframes: <a href=\"https://docs.microsoft.com/en-us/azure/machine-learning/team-data-science-process/spark-data-exploration-modeling\">https://docs.microsoft.com/en-us/azure/machine-learning/team-data-science-process/spark-data-exploration-modeling</a></li>\n</ol>\n\n<p>I have been able to load the dataframes into Pyspark and derive all my features. However I am facing memory limitations when trying to save or display few lines from the dataframes. </p>\n\n<p><code>WARN MemoryStore: Not enough space to cache rdd_1006_55 in memory! (computed 34.3 MB so far)</code>. </p>\n\n<p>I tried to increase the driver memory when launching Spark.</p>\n\n<p><code>pyspark --packages com.databricks:spark-csv_2.10:1.3.0 --master local[4] --driver-memory 4G</code></p>\n\n<p>It did not solved the issue.</p>\n\n<p><code>tests.repartition(1).write.format(\"com.databricks.spark.csv\").option(\"header\", \"true\").save(\"tests.csv\")</code> is too long and runs out of memory.</p>",
  "messages": [
    {
      "id": 310796,
      "postDate": "2018-04-08T16:11:19.533Z",
      "content": "<p>I am opening this topic for those who are interested in leveraging the full potential of TalkingData dataset with big data tools such as Apache Spark. </p>\n\n<p>Please could you share your tips?</p>\n\n<p>It can be for example the way you are partitioning the data to circumvent memory limitations, installation processes or any other tricks that can be useful for \"Spark Beginners\" to explore Talking Data full set and perhaps help them in doing some predictions. </p>\n\n<p>I have installed Pyspark 2.1.0 on a windows 10 computer (16 GB RAM and 4 cores) with anaconda Python 3.5 and Jupyter. Here I am listing the best resources that I have found so far:</p>\n\n<ol>\n<li>Installing Pyspark on windows: <a href=\"https://medium.com/@GalarnykMichael/install-spark-on-windows-pyspark-4498a5d8d66c\">https://medium.com/@GalarnykMichael/install-spark-on-windows-pyspark-4498a5d8d66c</a></li>\n<li>Using dataframes in Pyspark: <a href=\"https://www.analyticsvidhya.com/blog/2016/10/spark-dataframe-and-operations/\">https://www.analyticsvidhya.com/blog/2016/10/spark-dataframe-and-operations/</a></li>\n<li>A classification problem handle in Spark with summary statistics: <a href=\"https://mapr.com/blog/churn-prediction-pyspark-using-mllib-and-ml-packages/\">https://mapr.com/blog/churn-prediction-pyspark-using-mllib-and-ml-packages/</a></li>\n<li>Another machine learning example in Spark +statistics+ how to apply SQL queries on spark dataframes: <a href=\"https://docs.microsoft.com/en-us/azure/machine-learning/team-data-science-process/spark-data-exploration-modeling\">https://docs.microsoft.com/en-us/azure/machine-learning/team-data-science-process/spark-data-exploration-modeling</a></li>\n</ol>\n\n<p>I have been able to load the dataframes into Pyspark and derive all my features. However I am facing memory limitations when trying to save or display few lines from the dataframes. </p>\n\n<p><code>WARN MemoryStore: Not enough space to cache rdd_1006_55 in memory! (computed 34.3 MB so far)</code>. </p>\n\n<p>I tried to increase the driver memory when launching Spark.</p>\n\n<p><code>pyspark --packages com.databricks:spark-csv_2.10:1.3.0 --master local[4] --driver-memory 4G</code></p>\n\n<p>It did not solved the issue.</p>\n\n<p><code>tests.repartition(1).write.format(\"com.databricks.spark.csv\").option(\"header\", \"true\").save(\"tests.csv\")</code> is too long and runs out of memory.</p>",
      "rawMarkdown": "I am opening this topic for those who are interested in leveraging the full potential of TalkingData dataset with big data tools such as Apache Spark. \n\nPlease could you share your tips?\n\n It can be for example the way you are partitioning the data to circumvent memory limitations, installation processes or any other tricks that can be useful for \"Spark Beginners\" to explore Talking Data full set and perhaps help them in doing some predictions. \n\nI have installed Pyspark 2.1.0 on a windows 10 computer (16 GB RAM and 4 cores) with anaconda Python 3.5 and Jupyter. Here I am listing the best resources that I have found so far:\n\n 1. Installing Pyspark on windows: https://medium.com/@GalarnykMichael/install-spark-on-windows-pyspark-4498a5d8d66c\n 2. Using dataframes in Pyspark: https://www.analyticsvidhya.com/blog/2016/10/spark-dataframe-and-operations/\n 3. A classification problem handle in Spark with summary statistics: https://mapr.com/blog/churn-prediction-pyspark-using-mllib-and-ml-packages/\n 4.  Another machine learning example in Spark +statistics+ how to apply SQL queries on spark dataframes: https://docs.microsoft.com/en-us/azure/machine-learning/team-data-science-process/spark-data-exploration-modeling\n\nI have been able to load the dataframes into Pyspark and derive all my features. However I am facing memory limitations when trying to save or display few lines from the dataframes. \n\n` WARN MemoryStore: Not enough space to cache rdd_1006_55 in memory! (computed 34.3 MB so far)`. \n\nI tried to increase the driver memory when launching Spark.\n\n`pyspark --packages com.databricks:spark-csv_2.10:1.3.0 --master local[4] --driver-memory 4G`\n\nIt did not solved the issue.\n\n `tests.repartition(1).write.format(\"com.databricks.spark.csv\").option(\"header\", \"true\").save(\"tests.csv\")` is too long and runs out of memory.\n\n",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "310796": "I am opening this topic for those who are interested in leveraging the full potential of TalkingData dataset with big data tools such as Apache Spark. \n\nPlease could you share your tips?\n\n It can be for example the way you are partitioning the data to circumvent memory limitations, installation processes or any other tricks that can be useful for \"Spark Beginners\" to explore Talking Data full set and perhaps help them in doing some predictions. \n\nI have installed Pyspark 2.1.0 on a windows 10 computer (16 GB RAM and 4 cores) with anaconda Python 3.5 and Jupyter. Here I am listing the best resources that I have found so far:\n\n 1. Installing Pyspark on windows: https://medium.com/@GalarnykMichael/install-spark-on-windows-pyspark-4498a5d8d66c\n 2. Using dataframes in Pyspark: https://www.analyticsvidhya.com/blog/2016/10/spark-dataframe-and-operations/\n 3. A classification problem handle in Spark with summary statistics: https://mapr.com/blog/churn-prediction-pyspark-using-mllib-and-ml-packages/\n 4.  Another machine learning example in Spark +statistics+ how to apply SQL queries on spark dataframes: https://docs.microsoft.com/en-us/azure/machine-learning/team-data-science-process/spark-data-exploration-modeling\n\nI have been able to load the dataframes into Pyspark and derive all my features. However I am facing memory limitations when trying to save or display few lines from the dataframes. \n\n` WARN MemoryStore: Not enough space to cache rdd_1006_55 in memory! (computed 34.3 MB so far)`. \n\nI tried to increase the driver memory when launching Spark.\n\n`pyspark --packages com.databricks:spark-csv_2.10:1.3.0 --master local[4] --driver-memory 4G`\n\nIt did not solved the issue.\n\n `tests.repartition(1).write.format(\"com.databricks.spark.csv\").option(\"header\", \"true\").save(\"tests.csv\")` is too long and runs out of memory.\n\n"
  }
}