{
  "id": 545684,
  "title": "New to Kaggle and deep learning. Why am I getting this error while loading data?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/545684",
  "author_name": "",
  "post_date": "2024-11-11T15:40:21.711371100Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I am using the following code to load data.<br>\nimport pandas as pd<br>\nimport pyarrow.parquet as pq</p>\n<p>train_data_parts = []<br>\ntrain_path = '/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet'<br>\nfor i in range(10):<br>\n    part = pq.read_table(f\"{train_path}/partition_id={i}/part-0.parquet\").to_pandas()<br>\n    train_data_parts.append(part)</p>\n<p>train_data = pd.concat(train_data_parts, ignore_index=True)<br>\nprint(\"Training data shape:\", train_data.shape)<br>\nprint(train_data.head())</p>\n<p>It always runs out of memory. </p>",
  "messages": [
    {
      "id": "3042559",
      "postDate": "11/11/2024 15:40:21",
      "content": "<p>I am using the following code to load data.<br>\nimport pandas as pd<br>\nimport pyarrow.parquet as pq</p>\n<p>train_data_parts = []<br>\ntrain_path = '/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet'<br>\nfor i in range(10):<br>\n    part = pq.read_table(f\"{train_path}/partition_id={i}/part-0.parquet\").to_pandas()<br>\n    train_data_parts.append(part)</p>\n<p>train_data = pd.concat(train_data_parts, ignore_index=True)<br>\nprint(\"Training data shape:\", train_data.shape)<br>\nprint(train_data.head())</p>\n<p>It always runs out of memory. </p>",
      "rawMarkdown": "I am using the following code to load data.\nimport pandas as pd\nimport pyarrow.parquet as pq\n\ntrain_data_parts = []\ntrain_path = '/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet'\nfor i in range(10):\n    part = pq.read_table(f\"{train_path}/partition_id={i}/part-0.parquet\").to_pandas()\n    train_data_parts.append(part)\n\ntrain_data = pd.concat(train_data_parts, ignore_index=True)\nprint(\"Training data shape:\", train_data.shape)\nprint(train_data.head())\n\nIt always runs out of memory.",
      "votes": null
    },
    {
      "id": "3042630",
      "postDate": "11/11/2024 16:23:47",
      "content": "<p>There are millons of samples in the data given, so the kaggle kernel can be out of memory.You may need other computing power with more memory.</p>",
      "rawMarkdown": "There are millons of samples in the data given, so the kaggle kernel can be out of memory.You may need other computing power with more memory.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3042630,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "11/11/2024 16:23:47",
      "content": "<p>There are millons of samples in the data given, so the kaggle kernel can be out of memory.You may need other computing power with more memory.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3042559": "I am using the following code to load data.\nimport pandas as pd\nimport pyarrow.parquet as pq\n\ntrain_data_parts = []\ntrain_path = '/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet'\nfor i in range(10):\n    part = pq.read_table(f\"{train_path}/partition_id={i}/part-0.parquet\").to_pandas()\n    train_data_parts.append(part)\n\ntrain_data = pd.concat(train_data_parts, ignore_index=True)\nprint(\"Training data shape:\", train_data.shape)\nprint(train_data.head())\n\nIt always runs out of memory.",
    "3042630": "There are millons of samples in the data given, so the kaggle kernel can be out of memory.You may need other computing power with more memory."
  },
  "source": "meta"
}