{
  "id": 543328,
  "title": "fixing unexpected high ram usage?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543328",
  "author_name": "",
  "post_date": "2024-10-30T00:22:01.462781500Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I loaded a dataset using pandas and it has a size of 8 gb, the ram usage on kaggle's Draft Session metrics shows 9gb, then i execute the following cells: <br>\n<code>df = pl.from_pandas(train);\ndel train;\ngc.collect();\n</code> <br>\nthe ram usage after the execution is 24 gb, the gc.collect() outputs 0. <br>\n<code>lazy_df=df.lazy();\ndel df;\ngc.collect();</code> <br>\nram usage is still 24 gb and the gc.collect() outputs 0. <br>\nif i'm not wrong the memory should be free and it should be down to at most 2 or 1 gb.<br>\ncan anyone explain to me what's causing this high ram usage and how to fit it?</p>",
  "messages": [
    {
      "id": "3031634",
      "postDate": "10/30/2024 00:22:01",
      "content": "<p>I loaded a dataset using pandas and it has a size of 8 gb, the ram usage on kaggle's Draft Session metrics shows 9gb, then i execute the following cells: <br>\n<code>df = pl.from_pandas(train);\ndel train;\ngc.collect();\n</code> <br>\nthe ram usage after the execution is 24 gb, the gc.collect() outputs 0. <br>\n<code>lazy_df=df.lazy();\ndel df;\ngc.collect();</code> <br>\nram usage is still 24 gb and the gc.collect() outputs 0. <br>\nif i'm not wrong the memory should be free and it should be down to at most 2 or 1 gb.<br>\ncan anyone explain to me what's causing this high ram usage and how to fit it?</p>",
      "rawMarkdown": "I loaded a dataset using pandas and it has a size of 8 gb, the ram usage on kaggle's Draft Session metrics shows 9gb, then i execute the following cells: \n`df = pl.from_pandas(train);\ndel train;\ngc.collect();\n` \nthe ram usage after the execution is 24 gb, the gc.collect() outputs 0. \n`lazy_df=df.lazy();\ndel df;\ngc.collect();` \nram usage is still 24 gb and the gc.collect() outputs 0. \nif i'm not wrong the memory should be free and it should be down to at most 2 or 1 gb.\ncan anyone explain to me what's causing this high ram usage and how to fit it?",
      "votes": null
    },
    {
      "id": "3031883",
      "postDate": "10/30/2024 08:43:40",
      "content": "<p><a href=\"https://www.kaggle.com/younesbenalia\" target=\"_blank\">@younesbenalia</a> The results are being stored in IPython when you show the result of the dataframe at jupyter notebook. In Jupyter, for example, if it's processed at the 13th cell, the results remain in Out[13], _13, and _oh[13]. You need to clear these results.</p>\n<p>To clear them, instead of using del, you should assign the above Out[13] to a temporary variable and use the magic command %del to remove it (or use %reset out).</p>\n<p>please refer to <a href=\"https://www.kaggle.com/code/chumajin/memory-leakage-check\" target=\"_blank\">this notebook</a></p>",
      "rawMarkdown": "younesbenalia The results are being stored in IPython when you show the result of the dataframe at jupyter notebook. In Jupyter, for example, if it's processed at the 13th cell, the results remain in Out[13], _13, and _oh[13]. You need to clear these results.\n\nTo clear them, instead of using del, you should assign the above Out[13] to a temporary variable and use the magic command %del to remove it (or use %reset out).\n\nplease refer to [this notebook](https://www.kaggle.com/code/chumajin/memory-leakage-check)",
      "votes": null
    },
    {
      "id": "3032563",
      "postDate": "10/31/2024 04:59:54",
      "content": "<p>thank you for l this, this, however i'm not displaying any dataframes and my cells has no output to clear.<br>\nthe cause of this high ram usage is from the from_pandas function it's implementation makes it that both the pandas and polars df loads in memory.  I just avoided using it by using polars. btw what do you think about using fp16 to reduce the size of dataset ? will it affect the accuracy of the predictions ? <br>\none more thing to add, polars doesn't support fp16  :(</p>",
      "rawMarkdown": "thank you for l this, this, however i'm not displaying any dataframes and my cells has no output to clear.\nthe cause of this high ram usage is from the from_pandas function it's implementation makes it that both the pandas and polars df loads in memory.  I just avoided using it by using polars. btw what do you think about using fp16 to reduce the size of dataset ? will it affect the accuracy of the predictions ? \none more thing to add, polars doesn't support fp16  :(",
      "votes": null
    },
    {
      "id": "3032605",
      "postDate": "10/31/2024 06:27:32",
      "content": "<p><a href=\"https://www.kaggle.com/younesbenalia\" target=\"_blank\">@younesbenalia</a> OK. I see. I misunderstood. Float16 was used in the past AMEX competition.<br>\nThis is the great notebook. <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977</a></p>\n<p>However, as you mentioned, Polars doesn’t support fp16, so it might be a good idea to convert it to numpy or similar for evaluation!</p>",
      "rawMarkdown": "younesbenalia OK. I see. I misunderstood. Float16 was used in the past AMEX competition.\nThis is the great notebook. https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\n\nHowever, as you mentioned, Polars doesn’t support fp16, so it might be a good idea to convert it to numpy or similar for evaluation!",
      "votes": null
    },
    {
      "id": "3032631",
      "postDate": "10/31/2024 07:04:23",
      "content": "<p>Thank you for sharing!!</p>",
      "rawMarkdown": "Thank you for sharing!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3031883,
      "author_name": "chumajin",
      "author_url": "",
      "post_date": "10/30/2024 08:43:40",
      "content": "<p><a href=\"https://www.kaggle.com/younesbenalia\" target=\"_blank\">@younesbenalia</a> The results are being stored in IPython when you show the result of the dataframe at jupyter notebook. In Jupyter, for example, if it's processed at the 13th cell, the results remain in Out[13], _13, and _oh[13]. You need to clear these results.</p>\n<p>To clear them, instead of using del, you should assign the above Out[13] to a temporary variable and use the magic command %del to remove it (or use %reset out).</p>\n<p>please refer to <a href=\"https://www.kaggle.com/code/chumajin/memory-leakage-check\" target=\"_blank\">this notebook</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3032563,
          "author_name": "younesbenalia",
          "author_url": "",
          "post_date": "10/31/2024 04:59:54",
          "content": "<p>thank you for l this, this, however i'm not displaying any dataframes and my cells has no output to clear.<br>\nthe cause of this high ram usage is from the from_pandas function it's implementation makes it that both the pandas and polars df loads in memory.  I just avoided using it by using polars. btw what do you think about using fp16 to reduce the size of dataset ? will it affect the accuracy of the predictions ? <br>\none more thing to add, polars doesn't support fp16  :(</p>",
          "votes": null,
          "replies": [
            {
              "id": 3032605,
              "author_name": "chumajin",
              "author_url": "",
              "post_date": "10/31/2024 06:27:32",
              "content": "<p><a href=\"https://www.kaggle.com/younesbenalia\" target=\"_blank\">@younesbenalia</a> OK. I see. I misunderstood. Float16 was used in the past AMEX competition.<br>\nThis is the great notebook. <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977</a></p>\n<p>However, as you mentioned, Polars doesn’t support fp16, so it might be a good idea to convert it to numpy or similar for evaluation!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3032631,
                  "author_name": "younesbenalia",
                  "author_url": "",
                  "post_date": "10/31/2024 07:04:23",
                  "content": "<p>Thank you for sharing!!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3031634": "I loaded a dataset using pandas and it has a size of 8 gb, the ram usage on kaggle's Draft Session metrics shows 9gb, then i execute the following cells: \n`df = pl.from_pandas(train);\ndel train;\ngc.collect();\n` \nthe ram usage after the execution is 24 gb, the gc.collect() outputs 0. \n`lazy_df=df.lazy();\ndel df;\ngc.collect();` \nram usage is still 24 gb and the gc.collect() outputs 0. \nif i'm not wrong the memory should be free and it should be down to at most 2 or 1 gb.\ncan anyone explain to me what's causing this high ram usage and how to fit it?",
    "3031883": "younesbenalia The results are being stored in IPython when you show the result of the dataframe at jupyter notebook. In Jupyter, for example, if it's processed at the 13th cell, the results remain in Out[13], _13, and _oh[13]. You need to clear these results.\n\nTo clear them, instead of using del, you should assign the above Out[13] to a temporary variable and use the magic command %del to remove it (or use %reset out).\n\nplease refer to [this notebook](https://www.kaggle.com/code/chumajin/memory-leakage-check)",
    "3032563": "thank you for l this, this, however i'm not displaying any dataframes and my cells has no output to clear.\nthe cause of this high ram usage is from the from_pandas function it's implementation makes it that both the pandas and polars df loads in memory.  I just avoided using it by using polars. btw what do you think about using fp16 to reduce the size of dataset ? will it affect the accuracy of the predictions ? \none more thing to add, polars doesn't support fp16  :(",
    "3032605": "younesbenalia OK. I see. I misunderstood. Float16 was used in the past AMEX competition.\nThis is the great notebook. https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\n\nHowever, as you mentioned, Polars doesn’t support fp16, so it might be a good idea to convert it to numpy or similar for evaluation!",
    "3032631": "Thank you for sharing!!"
  },
  "source": "meta"
}