{
  "id": 544145,
  "title": "Solved | OOM | out of memory | How many RAM usage do we need? 128G is not even enough",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/544145",
  "author_name": "Mr RRR",
  "post_date": "2024-11-03T13:47:39.750000",
  "votes": 7,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi! May I ask how many RAM do you guys use? I use a 128G one but fail because it requires more memory. I am suspecting that if I do not do a good in memory management or this competition requires far more than that?</p>\n<p>Thanks for your sharing if possible, thanks!</p>",
  "messages": [
    {
      "id": 3035487,
      "postDate": "2024-11-03T13:47:39.750Z",
      "content": "<p>Hi! May I ask how many RAM do you guys use? I use a 128G one but fail because it requires more memory. I am suspecting that if I do not do a good in memory management or this competition requires far more than that?</p>\n<p>Thanks for your sharing if possible, thanks!</p>",
      "rawMarkdown": "Hi! May I ask how many RAM do you guys use? I use a 128G one but fail because it requires more memory. I am suspecting that if I do not do a good in memory management or this competition requires far more than that?\n\nThanks for your sharing if possible, thanks!",
      "votes": 7
    },
    {
      "id": 3035560,
      "postDate": "2024-11-03T15:37:10.093Z",
      "content": "<p>That is all you need!!!</p>\n<p>optimize your data typings and it using half of the memory.</p>\n<p>import pandas as pd<br>\nimport numpy as np</p>\n<p>def reduce_memory_usage(df):<br>\n    \"\"\"Reduce memory usage by optimizing data types.\"\"\"<br>\n    start_mem = df.memory_usage().sum() / 1024**2<br>\n    print(f\"Initial memory usage: {start_mem:.2f} MB\")</p>\n<pre><code>   df.:\n    col_type = df[].dtype\n\n     col_type != object:\n        c_min = df[].()\n        c_max = df[].()\n\n        # Convert int  to smallest possible  subtype\n         str(col_type)[:] == 'int':\n             c_min &gt; .iinfo(.int8).  c_max &lt; .iinfo(.int8).:\n                df[] = df[].astype(.int8)\n            elif c_min &gt; .iinfo(.int16).  c_max &lt; .iinfo(.int16).:\n                df[] = df[].astype(.int16)\n            elif c_min &gt; .iinfo(.int32).  c_max &lt; .iinfo(.int32).:\n                df[] = df[].astype(.int32)\n            elif c_min &gt; .iinfo(.int64).  c_max &lt; .iinfo(.int64).:\n                df[] = df[].astype(.int64)\n\n        # Convert   to smallest possible  subtype\n        :\n             c_min &gt; .finfo(.float16).  c_max &lt; .finfo(.float16).:\n                df[] = df[].astype(.float16)\n            elif c_min &gt; .finfo(.float32).  c_max &lt; .finfo(.float32).:\n                df[] = df[].astype(.float32)\n            :\n                df[] = df[].astype(.float64)\n\n    # Convert object  to category    are below threshold\n    :\n         df[].nunique() / len(df) &lt; :  # Adjust threshold as needed\n            df[] = df[].astype('category')\n\nend_mem = df.memory_usage().() / **\n(f)\n df\n</code></pre>",
      "rawMarkdown": "That is all you need!!!\n\noptimize your data typings and it using half of the memory.\n\nimport pandas as pd\nimport numpy as np\n\ndef reduce_memory_usage(df):\n    \"\"\"Reduce memory usage by optimizing data types.\"\"\"\n    start_mem = df.memory_usage().sum() / 1024**2\n    print(f\"Initial memory usage: {start_mem:.2f} MB\")\n\n    for col in df.columns:\n        col_type = df[col].dtype\n\n        if col_type != object:\n            c_min = df[col].min()\n            c_max = df[col].max()\n\n            # Convert int columns to smallest possible integer subtype\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)\n\n            # Convert float columns to smallest possible float subtype\n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n\n        # Convert object columns to category if unique values are below threshold\n        else:\n            if df[col].nunique() / len(df) < 0.5:  # Adjust threshold as needed\n                df[col] = df[col].astype('category')\n\n    end_mem = df.memory_usage().sum() / 1024**2\n    print(f\"Reduced memory usage: {end_mem:.2f} MB ({100 * (start_mem - end_mem) / start_mem:.1f}% reduction)\")\n    return df",
      "votes": 2,
      "replies": [
        {
          "id": 3035872,
          "postDate": "2024-11-04T00:22:05.433Z",
          "content": "<p>Hi! Farhan! Thanks! I have solved this problem using polars, however, my gpu out of memory….I am trying to fix this :) May I have your precious tips?</p>",
          "rawMarkdown": "Hi! Farhan! Thanks! I have solved this problem using polars, however, my gpu out of memory....I am trying to fix this :) May I have your precious tips?"
        },
        {
          "id": 3036077,
          "postDate": "2024-11-04T07:18:19.097Z",
          "content": "<p>After reading with Polars and converting to Pandas - almost all the columns are already in float32 (3 columns are integers) and using this code you are actually not converting anything lower than that because they don't fit in float16 and you are just losing time running this function, the only case it can help if you are creating additional features and want to convert them.</p>",
          "rawMarkdown": "After reading with Polars and converting to Pandas - almost all the columns are already in float32 (3 columns are integers) and using this code you are actually not converting anything lower than that because they don't fit in float16 and you are just losing time running this function, the only case it can help if you are creating additional features and want to convert them.",
          "replies": [
            {
              "id": 3036235,
              "postDate": "2024-11-04T11:39:34.720Z",
              "content": "<p>Yes, I agree with you after testing</p>",
              "rawMarkdown": "Yes, I agree with you after testing"
            }
          ]
        }
      ]
    },
    {
      "id": 3038360,
      "postDate": "2024-11-06T21:47:46.667Z",
      "content": "<p>You can basically do sequential learning. That's what I was doing.<br>\nJust load chunks of data (Rather than whole data) and partially train the model, then iterate over all the chunks and train model using those data.</p>",
      "rawMarkdown": "You can basically do sequential learning. That's what I was doing.\nJust load chunks of data (Rather than whole data) and partially train the model, then iterate over all the chunks and train model using those data.",
      "replies": [
        {
          "id": 3078869,
          "postDate": "2024-12-23T00:24:25.540Z",
          "content": "<p>I've been working with this but I'm having issues clearing the memory correctly after each batch. Any tips?</p>",
          "rawMarkdown": "I've been working with this but I'm having issues clearing the memory correctly after each batch. Any tips?"
        }
      ]
    },
    {
      "id": 3035514,
      "postDate": "2024-11-03T14:26:34.043Z",
      "content": "<p>Use lazy execution and prevent unnecessary copies of data. This will resolve the problem <a href=\"https://www.kaggle.com/larrylin666\" target=\"_blank\">@larrylin666</a> </p>",
      "rawMarkdown": "Use lazy execution and prevent unnecessary copies of data. This will resolve the problem @larrylin666 ",
      "replies": [
        {
          "id": 3035525,
          "postDate": "2024-11-03T14:43:41.223Z",
          "content": "<p>Hi! <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>\n<p>Thanks first! And actually I have used them, for example:</p>\n<p>scan = pl.scan_parquet(f\"{data_path}/train.parquet\").fetch(n_test)</p>\n<p>scan = pl.scan_parquet(f\"{data_path}/train.parquet\").collect()</p>\n<p>And delete the data which is not needed after preprocessing, however, I ran out of memory when training the model using CV strategy, may I ask what is the approximate peak of ram when you run your code?</p>",
          "rawMarkdown": "Hi! @ravi20076 \n\nThanks first! And actually I have used them, for example:\n\nscan = pl.scan_parquet(f\"{data_path}/train.parquet\").fetch(n_test)\n\nscan = pl.scan_parquet(f\"{data_path}/train.parquet\").collect()\n\nAnd delete the data which is not needed after preprocessing, however, I ran out of memory when training the model using CV strategy, may I ask what is the approximate peak of ram when you run your code?"
        },
        {
          "id": 3035873,
          "postDate": "2024-11-04T00:22:28.727Z",
          "content": "<p>Hi! Ravi! I have solved this problem using polars, however, my gpu out of memory….I am trying to fix this :) May I have your precious tips?</p>",
          "rawMarkdown": "Hi! Ravi! I have solved this problem using polars, however, my gpu out of memory....I am trying to fix this :) May I have your precious tips?"
        }
      ]
    },
    {
      "id": 3036524,
      "postDate": "2024-11-04T16:30:22.560Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 3036757,
          "postDate": "2024-11-04T22:28:52.450Z",
          "content": "<p>Hi! Thanks! I have solved this!</p>",
          "rawMarkdown": "Hi! Thanks! I have solved this!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3035560,
      "author_name": "Farhan Kardan",
      "author_url": "",
      "post_date": "2024-11-03T15:37:10.093000",
      "content": "<p>That is all you need!!!</p>\n<p>optimize your data typings and it using half of the memory.</p>\n<p>import pandas as pd<br>\nimport numpy as np</p>\n<p>def reduce_memory_usage(df):<br>\n    \"\"\"Reduce memory usage by optimizing data types.\"\"\"<br>\n    start_mem = df.memory_usage().sum() / 1024**2<br>\n    print(f\"Initial memory usage: {start_mem:.2f} MB\")</p>\n<pre><code>   df.:\n    col_type = df[].dtype\n\n     col_type != object:\n        c_min = df[].()\n        c_max = df[].()\n\n        # Convert int  to smallest possible  subtype\n         str(col_type)[:] == 'int':\n             c_min &gt; .iinfo(.int8).  c_max &lt; .iinfo(.int8).:\n                df[] = df[].astype(.int8)\n            elif c_min &gt; .iinfo(.int16).  c_max &lt; .iinfo(.int16).:\n                df[] = df[].astype(.int16)\n            elif c_min &gt; .iinfo(.int32).  c_max &lt; .iinfo(.int32).:\n                df[] = df[].astype(.int32)\n            elif c_min &gt; .iinfo(.int64).  c_max &lt; .iinfo(.int64).:\n                df[] = df[].astype(.int64)\n\n        # Convert   to smallest possible  subtype\n        :\n             c_min &gt; .finfo(.float16).  c_max &lt; .finfo(.float16).:\n                df[] = df[].astype(.float16)\n            elif c_min &gt; .finfo(.float32).  c_max &lt; .finfo(.float32).:\n                df[] = df[].astype(.float32)\n            :\n                df[] = df[].astype(.float64)\n\n    # Convert object  to category    are below threshold\n    :\n         df[].nunique() / len(df) &lt; :  # Adjust threshold as needed\n            df[] = df[].astype('category')\n\nend_mem = df.memory_usage().() / **\n(f)\n df\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 3035872,
          "author_name": "Mr RRR",
          "author_url": "",
          "post_date": "2024-11-04T00:22:05.433000",
          "content": "<p>Hi! Farhan! Thanks! I have solved this problem using polars, however, my gpu out of memory….I am trying to fix this :) May I have your precious tips?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3036077,
          "author_name": "Danu A.",
          "author_url": "",
          "post_date": "2024-11-04T07:18:19.097000",
          "content": "<p>After reading with Polars and converting to Pandas - almost all the columns are already in float32 (3 columns are integers) and using this code you are actually not converting anything lower than that because they don't fit in float16 and you are just losing time running this function, the only case it can help if you are creating additional features and want to convert them.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3036235,
              "author_name": "Mr RRR",
              "author_url": "",
              "post_date": "2024-11-04T11:39:34.720000",
              "content": "<p>Yes, I agree with you after testing</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3038360,
      "author_name": "Jay Shrivastava",
      "author_url": "",
      "post_date": "2024-11-06T21:47:46.667000",
      "content": "<p>You can basically do sequential learning. That's what I was doing.<br>\nJust load chunks of data (Rather than whole data) and partially train the model, then iterate over all the chunks and train model using those data.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3078869,
          "author_name": "Diego Delgado",
          "author_url": "",
          "post_date": "2024-12-23T00:24:25.540000",
          "content": "<p>I've been working with this but I'm having issues clearing the memory correctly after each batch. Any tips?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3035514,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2024-11-03T14:26:34.043000",
      "content": "<p>Use lazy execution and prevent unnecessary copies of data. This will resolve the problem <a href=\"https://www.kaggle.com/larrylin666\" target=\"_blank\">@larrylin666</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3035525,
          "author_name": "Mr RRR",
          "author_url": "",
          "post_date": "2024-11-03T14:43:41.223000",
          "content": "<p>Hi! <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>\n<p>Thanks first! And actually I have used them, for example:</p>\n<p>scan = pl.scan_parquet(f\"{data_path}/train.parquet\").fetch(n_test)</p>\n<p>scan = pl.scan_parquet(f\"{data_path}/train.parquet\").collect()</p>\n<p>And delete the data which is not needed after preprocessing, however, I ran out of memory when training the model using CV strategy, may I ask what is the approximate peak of ram when you run your code?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3035873,
          "author_name": "Mr RRR",
          "author_url": "",
          "post_date": "2024-11-04T00:22:28.727000",
          "content": "<p>Hi! Ravi! I have solved this problem using polars, however, my gpu out of memory….I am trying to fix this :) May I have your precious tips?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3036524,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-04T16:30:22.560000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3036757,
          "author_name": "Mr RRR",
          "author_url": "",
          "post_date": "2024-11-04T22:28:52.450000",
          "content": "<p>Hi! Thanks! I have solved this!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3035487": "Hi! May I ask how many RAM do you guys use? I use a 128G one but fail because it requires more memory. I am suspecting that if I do not do a good in memory management or this competition requires far more than that?\n\nThanks for your sharing if possible, thanks!",
    "3035560": "That is all you need!!!\n\noptimize your data typings and it using half of the memory.\n\nimport pandas as pd\nimport numpy as np\n\ndef reduce_memory_usage(df):\n    \"\"\"Reduce memory usage by optimizing data types.\"\"\"\n    start_mem = df.memory_usage().sum() / 1024**2\n    print(f\"Initial memory usage: {start_mem:.2f} MB\")\n\n    for col in df.columns:\n        col_type = df[col].dtype\n\n        if col_type != object:\n            c_min = df[col].min()\n            c_max = df[col].max()\n\n            # Convert int columns to smallest possible integer subtype\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)\n\n            # Convert float columns to smallest possible float subtype\n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n\n        # Convert object columns to category if unique values are below threshold\n        else:\n            if df[col].nunique() / len(df) < 0.5:  # Adjust threshold as needed\n                df[col] = df[col].astype('category')\n\n    end_mem = df.memory_usage().sum() / 1024**2\n    print(f\"Reduced memory usage: {end_mem:.2f} MB ({100 * (start_mem - end_mem) / start_mem:.1f}% reduction)\")\n    return df",
    "3038360": "You can basically do sequential learning. That's what I was doing.\nJust load chunks of data (Rather than whole data) and partially train the model, then iterate over all the chunks and train model using those data.",
    "3035514": "Use lazy execution and prevent unnecessary copies of data. This will resolve the problem @larrylin666 ",
    "3036524": ""
  }
}