{
  "id": 78394,
  "title": "RAM usage when loading data with pandas",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/78394",
  "author_name": "",
  "post_date": "2019-01-23T06:38:24.844344800Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>I think there is something wrong with the RAM used by Kaggle Kernel when calling pandas read_csv on big files.</p>\n\n<p>When I do on a Kernel:</p>\n\n<pre><code>train = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16}, usecols=['acoustic_data'])\n</code></pre>\n\n<p>RAM goes to 11 GB.</p>\n\n<p>But the DataFrame itself takes only 1.25 GB</p>\n\n<pre><code>&gt; train.memory_usage() / 1e9\nIndex            8.000000e-08\nacoustic_data    1.258291e+00\n</code></pre>\n\n<p>1.25 GB make sense since there is 600M row with 2 Bytes per row.</p>\n\n<p>If instead of int16 I load it using int32 the total RAM usage goes to 12.3, adding \"only\" the normal 1.25GB to RAM.</p>\n\n<p>When I am doing the same experimentation on a local environment I do not have the issue (1.4 GB only) it seems specific to Kaggle</p>\n\n<p>It is like the size of the file itself (train.csv is arround 9.3 GB) is added to RAM for an unknown reason.</p>",
  "messages": [
    {
      "id": "460190",
      "postDate": "01/23/2019 06:38:24",
      "content": "<p>Hello,</p>\n\n<p>I think there is something wrong with the RAM used by Kaggle Kernel when calling pandas read_csv on big files.</p>\n\n<p>When I do on a Kernel:</p>\n\n<pre><code>train = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16}, usecols=['acoustic_data'])\n</code></pre>\n\n<p>RAM goes to 11 GB.</p>\n\n<p>But the DataFrame itself takes only 1.25 GB</p>\n\n<pre><code>&gt; train.memory_usage() / 1e9\nIndex            8.000000e-08\nacoustic_data    1.258291e+00\n</code></pre>\n\n<p>1.25 GB make sense since there is 600M row with 2 Bytes per row.</p>\n\n<p>If instead of int16 I load it using int32 the total RAM usage goes to 12.3, adding \"only\" the normal 1.25GB to RAM.</p>\n\n<p>When I am doing the same experimentation on a local environment I do not have the issue (1.4 GB only) it seems specific to Kaggle</p>\n\n<p>It is like the size of the file itself (train.csv is arround 9.3 GB) is added to RAM for an unknown reason.</p>",
      "rawMarkdown": "Hello,\n\nI think there is something wrong with the RAM used by Kaggle Kernel when calling pandas read_csv on big files.\n\nWhen I do on a Kernel:\n\n    train = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16}, usecols=['acoustic_data'])\n\nRAM goes to 11 GB.\n\nBut the DataFrame itself takes only 1.25 GB\n\n    &gt; train.memory_usage() / 1e9\n    Index            8.000000e-08\n    acoustic_data    1.258291e+00\n\n\n1.25 GB make sense since there is 600M row with 2 Bytes per row.\n\nIf instead of int16 I load it using int32 the total RAM usage goes to 12.3, adding \"only\" the normal 1.25GB to RAM.\n\nWhen I am doing the same experimentation on a local environment I do not have the issue (1.4 GB only) it seems specific to Kaggle\n\nIt is like the size of the file itself (train.csv is arround 9.3 GB) is added to RAM for an unknown reason.",
      "votes": null
    },
    {
      "id": "464976",
      "postDate": "02/01/2019 23:09:33",
      "content": "<p>I preprocessed the training set into an gzip compressed hdf5 file, using int16 and float32, the size goes to 405M.</p>",
      "rawMarkdown": "I preprocessed the training set into an gzip compressed hdf5 file, using int16 and float32, the size goes to 405M.",
      "votes": null
    },
    {
      "id": "466193",
      "postDate": "02/04/2019 20:54:22",
      "content": "<p>I did the same with simple pickle of the DataFrame.</p>\n\n<p>But I am still wondering why Kaggle Notebook seems to keep the file in memory.</p>",
      "rawMarkdown": "I did the same with simple pickle of the DataFrame.\n\nBut I am still wondering why Kaggle Notebook seems to keep the file in memory.",
      "votes": null
    },
    {
      "id": "466279",
      "postDate": "02/05/2019 02:47:25",
      "content": "<p>The same happens when I load the data on a Kaggle Kernel. </p>",
      "rawMarkdown": "The same happens when I load the data on a Kaggle Kernel.",
      "votes": null
    },
    {
      "id": "467963",
      "postDate": "02/08/2019 03:31:27",
      "content": "<p>But when you load it to memory, it would still take up a lot of space.</p>",
      "rawMarkdown": "But when you load it to memory, it would still take up a lot of space.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 464976,
      "author_name": "kh40tika",
      "author_url": "",
      "post_date": "02/01/2019 23:09:33",
      "content": "<p>I preprocessed the training set into an gzip compressed hdf5 file, using int16 and float32, the size goes to 405M.</p>",
      "votes": null,
      "replies": [
        {
          "id": 467963,
          "author_name": "moushengxu",
          "author_url": "",
          "post_date": "02/08/2019 03:31:27",
          "content": "<p>But when you load it to memory, it would still take up a lot of space.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 466193,
      "author_name": "florenting",
      "author_url": "",
      "post_date": "02/04/2019 20:54:22",
      "content": "<p>I did the same with simple pickle of the DataFrame.</p>\n\n<p>But I am still wondering why Kaggle Notebook seems to keep the file in memory.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 466279,
      "author_name": "cdesilv",
      "author_url": "",
      "post_date": "02/05/2019 02:47:25",
      "content": "<p>The same happens when I load the data on a Kaggle Kernel. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "460190": "Hello,\n\nI think there is something wrong with the RAM used by Kaggle Kernel when calling pandas read_csv on big files.\n\nWhen I do on a Kernel:\n\n    train = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16}, usecols=['acoustic_data'])\n\nRAM goes to 11 GB.\n\nBut the DataFrame itself takes only 1.25 GB\n\n    &gt; train.memory_usage() / 1e9\n    Index            8.000000e-08\n    acoustic_data    1.258291e+00\n\n\n1.25 GB make sense since there is 600M row with 2 Bytes per row.\n\nIf instead of int16 I load it using int32 the total RAM usage goes to 12.3, adding \"only\" the normal 1.25GB to RAM.\n\nWhen I am doing the same experimentation on a local environment I do not have the issue (1.4 GB only) it seems specific to Kaggle\n\nIt is like the size of the file itself (train.csv is arround 9.3 GB) is added to RAM for an unknown reason.",
    "464976": "I preprocessed the training set into an gzip compressed hdf5 file, using int16 and float32, the size goes to 405M.",
    "466193": "I did the same with simple pickle of the DataFrame.\n\nBut I am still wondering why Kaggle Notebook seems to keep the file in memory.",
    "466279": "The same happens when I load the data on a Kaggle Kernel.",
    "467963": "But when you load it to memory, it would still take up a lot of space."
  },
  "source": "meta"
}