{
  "id": 368367,
  "title": "Has anyone tried PySpark?",
  "url": "/competitions/otto-recommender-system/discussion/368367",
  "author_name": "",
  "post_date": "2022-11-24T21:23:18.607323800Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>After a couple days struggling with the dataset I came to realize why not try to use PySpark (or any other Spark's flavours) on this competition. I have found myself trying to use some classic (Python-implemented) ML algorithms but I usually run out of memory or the compute lasts forever.  </p>\n<p>Has anyone given it a shot?</p>",
  "messages": [
    {
      "id": "2042668",
      "postDate": "11/24/2022 21:23:18",
      "content": "<p>After a couple days struggling with the dataset I came to realize why not try to use PySpark (or any other Spark's flavours) on this competition. I have found myself trying to use some classic (Python-implemented) ML algorithms but I usually run out of memory or the compute lasts forever.  </p>\n<p>Has anyone given it a shot?</p>",
      "rawMarkdown": "After a couple days struggling with the dataset I came to realize why not try to use PySpark (or any other Spark's flavours) on this competition. I have found myself trying to use some classic (Python-implemented) ML algorithms but I usually run out of memory or the compute lasts forever.  \n\nHas anyone given it a shot?",
      "votes": null
    },
    {
      "id": "2046554",
      "postDate": "11/28/2022 08:31:40",
      "content": "<p>I try to use pyspark in recall phase, and it can obtain the final recall results.</p>",
      "rawMarkdown": "I try to use pyspark in recall phase, and it can obtain the final recall results.",
      "votes": null
    },
    {
      "id": "3265401",
      "postDate": "08/07/2025 12:24:11",
      "content": "<p>Yes, it is possible to work in Pyspark here in Kaggle:<br>\n<a href=\"https://www.kaggle.com/code/beatafaron/pyspark-for-data-scientists-quick-shortcut\" target=\"_blank\">https://www.kaggle.com/code/beatafaron/pyspark-for-data-scientists-quick-shortcut</a></p>",
      "rawMarkdown": "Yes, it is possible to work in Pyspark here in Kaggle:\nhttps://www.kaggle.com/code/beatafaron/pyspark-for-data-scientists-quick-shortcut",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2046554,
      "author_name": "narutow",
      "author_url": "",
      "post_date": "11/28/2022 08:31:40",
      "content": "<p>I try to use pyspark in recall phase, and it can obtain the final recall results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3265401,
      "author_name": "beatafaron",
      "author_url": "",
      "post_date": "08/07/2025 12:24:11",
      "content": "<p>Yes, it is possible to work in Pyspark here in Kaggle:<br>\n<a href=\"https://www.kaggle.com/code/beatafaron/pyspark-for-data-scientists-quick-shortcut\" target=\"_blank\">https://www.kaggle.com/code/beatafaron/pyspark-for-data-scientists-quick-shortcut</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2042668": "After a couple days struggling with the dataset I came to realize why not try to use PySpark (or any other Spark's flavours) on this competition. I have found myself trying to use some classic (Python-implemented) ML algorithms but I usually run out of memory or the compute lasts forever.  \n\nHas anyone given it a shot?",
    "2046554": "I try to use pyspark in recall phase, and it can obtain the final recall results.",
    "3265401": "Yes, it is possible to work in Pyspark here in Kaggle:\nhttps://www.kaggle.com/code/beatafaron/pyspark-for-data-scientists-quick-shortcut"
  },
  "source": "meta"
}