{
  "id": 551897,
  "title": "Memory OOM issue when reading Parquet file on GPU (T4/P100), works only on TPU VM v3-8",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/551897",
  "author_name": "Hank.",
  "post_date": "2024-12-16T10:31:47.466000",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I am trying to read the train Parquet file using <code>pd.read_parquet()</code> in a Kaggle notebook. The command I use is:</p>\n<pre><code>df = pd.read_parquet()\n</code></pre>\n<p>However, I encounter an \"Out of Memory\" (OOM) error when running this on a GPU (either T4 or P100). Interestingly, the code works fine when I switch to a TPU VM v3-8. </p>\n<p>Has anyone experienced a similar issue when reading large Parquet files on GPUs? Is there any way to reduce memory usage or optimize the process to make it work on GPU instances (T4/P100)? Any suggestions or tips would be greatly appreciated!</p>\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": 3073561,
      "postDate": "2024-12-16T15:53:31.947Z",
      "content": "<p>This is an environment-specific RAM issue. TPU environments are equipped with over 300 GB of RAM, whereas GPU environments have less than 90 GB.</p>",
      "rawMarkdown": "This is an environment-specific RAM issue. TPU environments are equipped with over 300 GB of RAM, whereas GPU environments have less than 90 GB.",
      "votes": 2,
      "replies": [
        {
          "id": 3073581,
          "postDate": "2024-12-16T16:06:58.703Z",
          "content": "<p>Thanks. So how to load the entire data in GPU environment, as the competition does not allow to use TPU</p>",
          "rawMarkdown": "Thanks. So how to load the entire data in GPU environment, as the competition does not allow to use TPU",
          "replies": [
            {
              "id": 3073692,
              "postDate": "2024-12-16T18:35:54.253Z",
              "content": "<p>I am also eager to know this. I want to use all of the data, but how?</p>",
              "rawMarkdown": "I am also eager to know this. I want to use all of the data, but how?"
            },
            {
              "id": 3073786,
              "postDate": "2024-12-16T23:25:20.290Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3073787,
              "postDate": "2024-12-16T23:26:18.167Z",
              "content": "<p>def load_data(file_paths):<br>\n    return pd.concat([pd.read_parquet(file) for file in file_paths])<br>\ntrain_files = [f'/your_path/train.parquet/partition_id={i}/part-0.parquet' for i in range(10)]</p>\n<p>train_df = load_data(train_files[0:10])</p>",
              "rawMarkdown": "def load_data(file_paths):\n    return pd.concat([pd.read_parquet(file) for file in file_paths])\ntrain_files = [f'/your_path/train.parquet/partition_id={i}/part-0.parquet' for i in range(10)]\n\ntrain_df = load_data(train_files[0:10])"
            },
            {
              "id": 3073888,
              "postDate": "2024-12-17T04:21:55.303Z",
              "content": "<p>Still face \"Your notebook tried to allocate more memory than is available. It has restarted.\". I've tried both T4 and P100</p>",
              "rawMarkdown": "Still face \"Your notebook tried to allocate more memory than is available. It has restarted.\". I've tried both T4 and P100"
            }
          ]
        }
      ]
    },
    {
      "id": 3073312,
      "postDate": "2024-12-16T10:31:47.467Z",
      "content": "<p>Hi everyone,</p>\n<p>I am trying to read the train Parquet file using <code>pd.read_parquet()</code> in a Kaggle notebook. The command I use is:</p>\n<pre><code>df = pd.read_parquet()\n</code></pre>\n<p>However, I encounter an \"Out of Memory\" (OOM) error when running this on a GPU (either T4 or P100). Interestingly, the code works fine when I switch to a TPU VM v3-8. </p>\n<p>Has anyone experienced a similar issue when reading large Parquet files on GPUs? Is there any way to reduce memory usage or optimize the process to make it work on GPU instances (T4/P100)? Any suggestions or tips would be greatly appreciated!</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Hi everyone,\n\nI am trying to read the train Parquet file using `pd.read_parquet()` in a Kaggle notebook. The command I use is:\n\n```python\ndf = pd.read_parquet(f'{input_path}/train.parquet')\n```\n\nHowever, I encounter an \"Out of Memory\" (OOM) error when running this on a GPU (either T4 or P100). Interestingly, the code works fine when I switch to a TPU VM v3-8. \n\nHas anyone experienced a similar issue when reading large Parquet files on GPUs? Is there any way to reduce memory usage or optimize the process to make it work on GPU instances (T4/P100)? Any suggestions or tips would be greatly appreciated!\n\nThanks in advance!",
      "votes": 2
    },
    {
      "id": 3076626,
      "postDate": "2024-12-20T06:02:56.980Z",
      "content": "<p>same thing.<br>\nCan't retrain XGB outside of the Kaggle</p>",
      "rawMarkdown": "same thing.\nCan't retrain XGB outside of the Kaggle"
    }
  ],
  "comments": [
    {
      "id": 3073561,
      "author_name": "Moujahid bou-laouz",
      "author_url": "",
      "post_date": "2024-12-16T15:53:31.947000",
      "content": "<p>This is an environment-specific RAM issue. TPU environments are equipped with over 300 GB of RAM, whereas GPU environments have less than 90 GB.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3073581,
          "author_name": "Hank.",
          "author_url": "",
          "post_date": "2024-12-16T16:06:58.703000",
          "content": "<p>Thanks. So how to load the entire data in GPU environment, as the competition does not allow to use TPU</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3073692,
              "author_name": "小亨瑞",
              "author_url": "",
              "post_date": "2024-12-16T18:35:54.253000",
              "content": "<p>I am also eager to know this. I want to use all of the data, but how?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3073786,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-12-16T23:25:20.290000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3073787,
              "author_name": "Moujahid bou-laouz",
              "author_url": "",
              "post_date": "2024-12-16T23:26:18.167000",
              "content": "<p>def load_data(file_paths):<br>\n    return pd.concat([pd.read_parquet(file) for file in file_paths])<br>\ntrain_files = [f'/your_path/train.parquet/partition_id={i}/part-0.parquet' for i in range(10)]</p>\n<p>train_df = load_data(train_files[0:10])</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3073888,
              "author_name": "Hank.",
              "author_url": "",
              "post_date": "2024-12-17T04:21:55.303000",
              "content": "<p>Still face \"Your notebook tried to allocate more memory than is available. It has restarted.\". I've tried both T4 and P100</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3076626,
      "author_name": "M4XD",
      "author_url": "",
      "post_date": "2024-12-20T06:02:56.980000",
      "content": "<p>same thing.<br>\nCan't retrain XGB outside of the Kaggle</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3073561": "This is an environment-specific RAM issue. TPU environments are equipped with over 300 GB of RAM, whereas GPU environments have less than 90 GB.",
    "3073312": "Hi everyone,\n\nI am trying to read the train Parquet file using `pd.read_parquet()` in a Kaggle notebook. The command I use is:\n\n```python\ndf = pd.read_parquet(f'{input_path}/train.parquet')\n```\n\nHowever, I encounter an \"Out of Memory\" (OOM) error when running this on a GPU (either T4 or P100). Interestingly, the code works fine when I switch to a TPU VM v3-8. \n\nHas anyone experienced a similar issue when reading large Parquet files on GPUs? Is there any way to reduce memory usage or optimize the process to make it work on GPU instances (T4/P100)? Any suggestions or tips would be greatly appreciated!\n\nThanks in advance!",
    "3076626": "same thing.\nCan't retrain XGB outside of the Kaggle"
  }
}