{
  "id": 373257,
  "title": "PySpark: How to efficiently read the data & write as parquet",
  "url": "/competitions/otto-recommender-system/discussion/373257",
  "author_name": "",
  "post_date": "2022-12-20T13:27:27.784686600Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have created the following PySpark script, which reads-in the JSON files, converts them to a tabular PySpark dataframe and outputs the files as parquet. </p>\n<p>Using Spark, this code is able to run with a simple Kaggle kernel (in about an hour) w/o running into memory issues.</p>\n<p><a href=\"https://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet\" target=\"_blank\">https://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet</a></p>",
  "messages": [
    {
      "id": "2070927",
      "postDate": "12/20/2022 13:27:27",
      "content": "<p>I have created the following PySpark script, which reads-in the JSON files, converts them to a tabular PySpark dataframe and outputs the files as parquet. </p>\n<p>Using Spark, this code is able to run with a simple Kaggle kernel (in about an hour) w/o running into memory issues.</p>\n<p><a href=\"https://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet\" target=\"_blank\">https://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet</a></p>",
      "rawMarkdown": "I have created the following PySpark script, which reads-in the JSON files, converts them to a tabular PySpark dataframe and outputs the files as parquet. \n\nUsing Spark, this code is able to run with a simple Kaggle kernel (in about an hour) w/o running into memory issues.\n\nhttps://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet",
      "votes": null
    },
    {
      "id": "2163841",
      "postDate": "03/01/2023 05:44:01",
      "content": "<p>Hi. <a href=\"https://www.kaggle.com/ikogias\" target=\"_blank\">@ikogias</a>  </p>\n<p>Thanks for sharing. I'm also interested to process it with Spark, but I'm not that familiar with spark.\\</p>\n<p>Just one more question, it seems that you notebook was ran on kaggle notebook where there is no extra resource applied. Is there benifit to run it with spark without extra resource allocated? should we set clusters to run the spark code？</p>",
      "rawMarkdown": "Hi. @ikogias  \n\nThanks for sharing. I'm also interested to process it with Spark, but I'm not that familiar with spark.\\\n\nJust one more question, it seems that you notebook was ran on kaggle notebook where there is no extra resource applied. Is there benifit to run it with spark without extra resource allocated? should we set clusters to run the spark code？",
      "votes": null
    },
    {
      "id": "2251206",
      "postDate": "05/09/2023 06:56:22",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/huaguo\" target=\"_blank\">@huaguo</a>, apologies for the late reply! To run this Spark code you don't need to set up any clusters or allocate any extra resource. Yes, there is a benefit in running Spark even without any clusters, as it manages to parallelise computations even when using a single computing instance. See here for an explanation:<br>\n<a href=\"https://www.databricks.com/blog/2018/05/03/benchmarking-apache-spark-on-a-single-node-machine.html\" target=\"_blank\">Benchmarking Apache Spark on a Single Node Machine</a></p>",
      "rawMarkdown": "Hey @huaguo, apologies for the late reply! To run this Spark code you don't need to set up any clusters or allocate any extra resource. Yes, there is a benefit in running Spark even without any clusters, as it manages to parallelise computations even when using a single computing instance. See here for an explanation:\n[Benchmarking Apache Spark on a Single Node Machine](https://www.databricks.com/blog/2018/05/03/benchmarking-apache-spark-on-a-single-node-machine.html)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2163841,
      "author_name": "huaguo",
      "author_url": "",
      "post_date": "03/01/2023 05:44:01",
      "content": "<p>Hi. <a href=\"https://www.kaggle.com/ikogias\" target=\"_blank\">@ikogias</a>  </p>\n<p>Thanks for sharing. I'm also interested to process it with Spark, but I'm not that familiar with spark.\\</p>\n<p>Just one more question, it seems that you notebook was ran on kaggle notebook where there is no extra resource applied. Is there benifit to run it with spark without extra resource allocated? should we set clusters to run the spark code？</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2251206,
      "author_name": "ikogias",
      "author_url": "",
      "post_date": "05/09/2023 06:56:22",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/huaguo\" target=\"_blank\">@huaguo</a>, apologies for the late reply! To run this Spark code you don't need to set up any clusters or allocate any extra resource. Yes, there is a benefit in running Spark even without any clusters, as it manages to parallelise computations even when using a single computing instance. See here for an explanation:<br>\n<a href=\"https://www.databricks.com/blog/2018/05/03/benchmarking-apache-spark-on-a-single-node-machine.html\" target=\"_blank\">Benchmarking Apache Spark on a Single Node Machine</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2070927": "I have created the following PySpark script, which reads-in the JSON files, converts them to a tabular PySpark dataframe and outputs the files as parquet. \n\nUsing Spark, this code is able to run with a simple Kaggle kernel (in about an hour) w/o running into memory issues.\n\nhttps://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet",
    "2163841": "Hi. @ikogias  \n\nThanks for sharing. I'm also interested to process it with Spark, but I'm not that familiar with spark.\\\n\nJust one more question, it seems that you notebook was ran on kaggle notebook where there is no extra resource applied. Is there benifit to run it with spark without extra resource allocated? should we set clusters to run the spark code？",
    "2251206": "Hey @huaguo, apologies for the late reply! To run this Spark code you don't need to set up any clusters or allocate any extra resource. Yes, there is a benefit in running Spark even without any clusters, as it manages to parallelise computations even when using a single computing instance. See here for an explanation:\n[Benchmarking Apache Spark on a Single Node Machine](https://www.databricks.com/blog/2018/05/03/benchmarking-apache-spark-on-a-single-node-machine.html)"
  },
  "source": "meta"
}