{
  "id": 380515,
  "title": "\"Your notebook tried to allocate more memory than is available. It has restarted.\"",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/380515",
  "author_name": "Arvind Devarkonda",
  "post_date": "2023-01-23T12:47:16.934000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>i don't know i am only getting this error i saw 2-3 kernels in which they are too reading file using <code>pd.read_parquet()</code> and displaying dataframe but they didn't got this error</p>",
  "messages": [
    {
      "id": 2112547,
      "postDate": "2023-01-23T17:20:53.440Z",
      "content": "<p>I'm having the same issues using the provided kaggle notebook environment. What's helped is reading only a subset of a file at a time. Here's some code that works for me: </p>\n<pre><code>import pyarrow.parquet as pq\n\nfile = pq.ParquetFile(datapath/\"train_meta.parquet\", buffer_size=10_000)\nbatches = file.iter_batches(batch_size=10_000)\ncurr_df = next(batches).to_pandas()\ncurr_df.head()\n</code></pre>",
      "rawMarkdown": "I'm having the same issues using the provided kaggle notebook environment. What's helped is reading only a subset of a file at a time. Here's some code that works for me: \n\n```\nimport pyarrow.parquet as pq\n\nfile = pq.ParquetFile(datapath/\"train_meta.parquet\", buffer_size=10_000)\nbatches = file.iter_batches(batch_size=10_000)\ncurr_df = next(batches).to_pandas()\ncurr_df.head()\n```",
      "votes": 5
    },
    {
      "id": 2112161,
      "postDate": "2023-01-23T12:47:16.933Z",
      "content": "<p>i don't know i am only getting this error i saw 2-3 kernels in which they are too reading file using <code>pd.read_parquet()</code> and displaying dataframe but they didn't got this error</p>",
      "rawMarkdown": "i don't know i am only getting this error i saw 2-3 kernels in which they are too reading file using ```pd.read_parquet()``` and displaying dataframe but they didn't got this error",
      "votes": 1
    },
    {
      "id": 2140618,
      "postDate": "2023-02-11T23:18:52.727Z",
      "content": "<p>I was getting out of memory errors randomly when calling <code>head()</code> or <code>shape</code> attribute on <em>meta train table</em> or sampled tables that occupied &lt;4gb of RAM. All of a sudden memory consumption was getting up above 30 gb. I had a hard time to predict what was the cause as this seemed to appear randomly. Not sure if its Kaggle environment but on Google Colab I do not have such issues. </p>",
      "rawMarkdown": "I was getting out of memory errors randomly when calling `head()` or `shape` attribute on *meta train table* or sampled tables that occupied <4gb of RAM. All of a sudden memory consumption was getting up above 30 gb. I had a hard time to predict what was the cause as this seemed to appear randomly. Not sure if its Kaggle environment but on Google Colab I do not have such issues. "
    },
    {
      "id": 2112167,
      "postDate": "2023-01-23T12:51:24.460Z",
      "content": "<p>Do you have a GPU attached?  If so, you're running on an instance with less main system RAM and are likely to see that error.</p>\n<p>If that's the problem, you can \"fix\" it by turning off the GPU accelerator.  I don't know if anybody has managed to load up the training metadata when there's a GPU attached. There are ~131M rows in that file and loading it seems to consume ~12GB RAM.</p>",
      "rawMarkdown": "Do you have a GPU attached?  If so, you're running on an instance with less main system RAM and are likely to see that error.\n\nIf that's the problem, you can \"fix\" it by turning off the GPU accelerator.  I don't know if anybody has managed to load up the training metadata when there's a GPU attached. There are ~131M rows in that file and loading it seems to consume ~12GB RAM.",
      "replies": [
        {
          "id": 2112271,
          "postDate": "2023-01-23T14:02:46.197Z",
          "content": "<p>i am not using gpu </p>",
          "rawMarkdown": "i am not using gpu "
        }
      ]
    },
    {
      "id": 2112449,
      "postDate": "2023-01-23T16:15:58.103Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2112547,
      "author_name": "stepheniota",
      "author_url": "",
      "post_date": "2023-01-23T17:20:53.440000",
      "content": "<p>I'm having the same issues using the provided kaggle notebook environment. What's helped is reading only a subset of a file at a time. Here's some code that works for me: </p>\n<pre><code>import pyarrow.parquet as pq\n\nfile = pq.ParquetFile(datapath/\"train_meta.parquet\", buffer_size=10_000)\nbatches = file.iter_batches(batch_size=10_000)\ncurr_df = next(batches).to_pandas()\ncurr_df.head()\n</code></pre>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2140618,
      "author_name": "RafalP",
      "author_url": "",
      "post_date": "2023-02-11T23:18:52.727000",
      "content": "<p>I was getting out of memory errors randomly when calling <code>head()</code> or <code>shape</code> attribute on <em>meta train table</em> or sampled tables that occupied &lt;4gb of RAM. All of a sudden memory consumption was getting up above 30 gb. I had a hard time to predict what was the cause as this seemed to appear randomly. Not sure if its Kaggle environment but on Google Colab I do not have such issues. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2112167,
      "author_name": "Andrew",
      "author_url": "",
      "post_date": "2023-01-23T12:51:24.460000",
      "content": "<p>Do you have a GPU attached?  If so, you're running on an instance with less main system RAM and are likely to see that error.</p>\n<p>If that's the problem, you can \"fix\" it by turning off the GPU accelerator.  I don't know if anybody has managed to load up the training metadata when there's a GPU attached. There are ~131M rows in that file and loading it seems to consume ~12GB RAM.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2112271,
          "author_name": "Arvind Devarkonda",
          "author_url": "",
          "post_date": "2023-01-23T14:02:46.197000",
          "content": "<p>i am not using gpu </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2112449,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-23T16:15:58.103000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2112547": "I'm having the same issues using the provided kaggle notebook environment. What's helped is reading only a subset of a file at a time. Here's some code that works for me: \n\n```\nimport pyarrow.parquet as pq\n\nfile = pq.ParquetFile(datapath/\"train_meta.parquet\", buffer_size=10_000)\nbatches = file.iter_batches(batch_size=10_000)\ncurr_df = next(batches).to_pandas()\ncurr_df.head()\n```",
    "2112161": "i don't know i am only getting this error i saw 2-3 kernels in which they are too reading file using ```pd.read_parquet()``` and displaying dataframe but they didn't got this error",
    "2140618": "I was getting out of memory errors randomly when calling `head()` or `shape` attribute on *meta train table* or sampled tables that occupied <4gb of RAM. All of a sudden memory consumption was getting up above 30 gb. I had a hard time to predict what was the cause as this seemed to appear randomly. Not sure if its Kaggle environment but on Google Colab I do not have such issues. ",
    "2112167": "Do you have a GPU attached?  If so, you're running on an instance with less main system RAM and are likely to see that error.\n\nIf that's the problem, you can \"fix\" it by turning off the GPU accelerator.  I don't know if anybody has managed to load up the training metadata when there's a GPU attached. There are ~131M rows in that file and loading it seems to consume ~12GB RAM.",
    "2112449": ""
  }
}