{
  "id": 332009,
  "title": "My Approach with Pyspark",
  "url": "/competitions/amex-default-prediction/discussion/332009",
  "author_name": "",
  "post_date": "2022-06-19T22:57:02.093000600Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Here are my notebooks to tackle this competition with pyspark. I know that this is not a best solution, conversely we can learn together and maybe know the way to optimize this approach more. The notebooks are:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/rakkaalhazimi/export-large-dataset-to-spark\" target=\"_blank\">Export Large Dataset into Pyspark</a></li>\n<li><a href=\"https://www.kaggle.com/code/rakkaalhazimi/ml-with-pyspark/notebook\" target=\"_blank\">ML with Pyspark (weak baseline)</a></li>\n</ol>\n<p>As for the EDA I might finish that with pyspark, yet I will take a look at other Kagglers notebook instead.</p>\n<p>Feel free if you want to optimize my notebooks or just take a brief look into it.<br>\nCheers :D</p>",
  "messages": [
    {
      "id": "1825995",
      "postDate": "06/19/2022 22:57:02",
      "content": "<p>Here are my notebooks to tackle this competition with pyspark. I know that this is not a best solution, conversely we can learn together and maybe know the way to optimize this approach more. The notebooks are:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/rakkaalhazimi/export-large-dataset-to-spark\" target=\"_blank\">Export Large Dataset into Pyspark</a></li>\n<li><a href=\"https://www.kaggle.com/code/rakkaalhazimi/ml-with-pyspark/notebook\" target=\"_blank\">ML with Pyspark (weak baseline)</a></li>\n</ol>\n<p>As for the EDA I might finish that with pyspark, yet I will take a look at other Kagglers notebook instead.</p>\n<p>Feel free if you want to optimize my notebooks or just take a brief look into it.<br>\nCheers :D</p>",
      "rawMarkdown": "Here are my notebooks to tackle this competition with pyspark. I know that this is not a best solution, conversely we can learn together and maybe know the way to optimize this approach more. The notebooks are:\n\n1. [Export Large Dataset into Pyspark](https://www.kaggle.com/code/rakkaalhazimi/export-large-dataset-to-spark)\n2. [ML with Pyspark (weak baseline)](https://www.kaggle.com/code/rakkaalhazimi/ml-with-pyspark/notebook)\n\nAs for the EDA I might finish that with pyspark, yet I will take a look at other Kagglers notebook instead.\n\nFeel free if you want to optimize my notebooks or just take a brief look into it.\nCheers :D",
      "votes": null
    },
    {
      "id": "1826046",
      "postDate": "06/20/2022 01:29:49",
      "content": "<p>Good approach!</p>",
      "rawMarkdown": "Good approach!",
      "votes": null
    },
    {
      "id": "1827986",
      "postDate": "06/21/2022 12:38:07",
      "content": "<p>I have done some studies with pyspark in a past <a href=\"https://www.kaggle.com/code/paulojunqueira/hands-on-pyspark-introduction-101\" target=\"_blank\">notebook</a>, just to practice…but I wonder how pyspark handles in the Kaggle environment as it is made for distributed systems… Anyways, nice work!</p>",
      "rawMarkdown": "I have done some studies with pyspark in a past [notebook](https://www.kaggle.com/code/paulojunqueira/hands-on-pyspark-introduction-101), just to practice...but I wonder how pyspark handles in the Kaggle environment as it is made for distributed systems... Anyways, nice work!",
      "votes": null
    },
    {
      "id": "1830305",
      "postDate": "06/23/2022 10:29:58",
      "content": "<p>What a good notebook, how can I overlooked it. Thank you for sharing.</p>",
      "rawMarkdown": "What a good notebook, how can I overlooked it. Thank you for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1826046,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "06/20/2022 01:29:49",
      "content": "<p>Good approach!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1827986,
      "author_name": "paulojunqueira",
      "author_url": "",
      "post_date": "06/21/2022 12:38:07",
      "content": "<p>I have done some studies with pyspark in a past <a href=\"https://www.kaggle.com/code/paulojunqueira/hands-on-pyspark-introduction-101\" target=\"_blank\">notebook</a>, just to practice…but I wonder how pyspark handles in the Kaggle environment as it is made for distributed systems… Anyways, nice work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1830305,
          "author_name": "rakkaalhazimi",
          "author_url": "",
          "post_date": "06/23/2022 10:29:58",
          "content": "<p>What a good notebook, how can I overlooked it. Thank you for sharing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1825995": "Here are my notebooks to tackle this competition with pyspark. I know that this is not a best solution, conversely we can learn together and maybe know the way to optimize this approach more. The notebooks are:\n\n1. [Export Large Dataset into Pyspark](https://www.kaggle.com/code/rakkaalhazimi/export-large-dataset-to-spark)\n2. [ML with Pyspark (weak baseline)](https://www.kaggle.com/code/rakkaalhazimi/ml-with-pyspark/notebook)\n\nAs for the EDA I might finish that with pyspark, yet I will take a look at other Kagglers notebook instead.\n\nFeel free if you want to optimize my notebooks or just take a brief look into it.\nCheers :D",
    "1826046": "Good approach!",
    "1827986": "I have done some studies with pyspark in a past [notebook](https://www.kaggle.com/code/paulojunqueira/hands-on-pyspark-introduction-101), just to practice...but I wonder how pyspark handles in the Kaggle environment as it is made for distributed systems... Anyways, nice work!",
    "1830305": "What a good notebook, how can I overlooked it. Thank you for sharing."
  },
  "source": "meta"
}