{
  "id": 21561,
  "title": "tooooo much data on test, use Spark?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21561",
  "author_name": "",
  "post_date": "2016-06-09T22:31:38.550Z",
  "votes": null,
  "comment_count": 1,
  "views": 662,
  "content": "<p>Hello:\nI am new to Kaggle and this is my first try in kaggle. I found it too much data about 3 million in testing data set, which takes me a huge time to run my code. Should I use Spark? or is it possible to run through jupyter.\nThank you </p>",
  "messages": [
    {
      "id": "123134",
      "postDate": "06/09/2016 22:31:38",
      "content": "<p>Hello:\nI am new to Kaggle and this is my first try in kaggle. I found it too much data about 3 million in testing data set, which takes me a huge time to run my code. Should I use Spark? or is it possible to run through jupyter.\nThank you </p>",
      "rawMarkdown": "Hello:\r\nI am new to Kaggle and this is my first try in kaggle. I found it too much data about 3 million in testing data set, which takes me a huge time to run my code. Should I use Spark? or is it possible to run through jupyter.\r\nThank you",
      "votes": null
    },
    {
      "id": "123177",
      "postDate": "06/10/2016 05:26:48",
      "content": "<p><a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20896/leakage-solution-with-spark-sql-pyspark/122256\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20896/leakage-solution-with-spark-sql-pyspark/122256</a></p>",
      "rawMarkdown": "https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20896/leakage-solution-with-spark-sql-pyspark/122256",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 123177,
      "author_name": "vykhand",
      "author_url": "",
      "post_date": "06/10/2016 05:26:48",
      "content": "<p><a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20896/leakage-solution-with-spark-sql-pyspark/122256\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20896/leakage-solution-with-spark-sql-pyspark/122256</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "123134": "Hello:\r\nI am new to Kaggle and this is my first try in kaggle. I found it too much data about 3 million in testing data set, which takes me a huge time to run my code. Should I use Spark? or is it possible to run through jupyter.\r\nThank you",
    "123177": "https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20896/leakage-solution-with-spark-sql-pyspark/122256"
  },
  "source": "meta"
}