{
  "id": 580485,
  "title": "How to deal with memory issues",
  "url": "/competitions/drw-crypto-market-prediction/discussion/580485",
  "author_name": "",
  "post_date": "2025-05-24T09:42:57.961549300Z",
  "votes": 40,
  "comment_count": 14,
  "views": 0,
  "content": "<p>The dataset for this competition is very large, both in terms of rows and columns, and it's likely to run into memory issues during EDA or model training. There are many ways to handle such challenges, but one simple  method is to reduce the memory footprint of the dataset. This can be done by changing data types to lower-precision formats, which helps save memory space. Below is a code snippet that I frequently use when working with large datasets on Kaggle.</p>\n<p>Hope you find it useful, and best of luck with the competition!</p>\n<pre><code> ():    \n    (, dataset)\n    initial_mem_usage = dataframe.memory_usage().() / **\n\n     col  dataframe.columns:\n        col_type = dataframe[col].dtype\n\n        c_min = dataframe[col].()\n        c_max = dataframe[col].()\n         (col_type)[:] == :\n             c_min &gt; np.iinfo(np.int8).  c_max &lt; np.iinfo(np.int8).:\n                dataframe[col] = dataframe[col].astype(np.int8)\n             c_min &gt; np.iinfo(np.int16).  c_max &lt; np.iinfo(np.int16).:\n                dataframe[col] = dataframe[col].astype(np.int16)\n             c_min &gt; np.iinfo(np.int32).  c_max &lt; np.iinfo(np.int32).:\n                dataframe[col] = dataframe[col].astype(np.int32)\n             c_min &gt; np.iinfo(np.int64).  c_max &lt; np.iinfo(np.int64).:\n                dataframe[col] = dataframe[col].astype(np.int64)\n        :\n             c_min &gt; np.finfo(np.float16).  c_max &lt; np.finfo(np.float16).:\n                dataframe[col] = dataframe[col].astype(np.float16)\n             c_min &gt; np.finfo(np.float32).  c_max &lt; np.finfo(np.float32).:\n                dataframe[col] = dataframe[col].astype(np.float32)\n            :\n                dataframe[col] = dataframe[col].astype(np.float64)\n\n    final_mem_usage = dataframe.memory_usage().() / **\n    (.(initial_mem_usage))\n    (.(final_mem_usage))\n    (.( * (initial_mem_usage - final_mem_usage) / initial_mem_usage))\n\n     dataframe\n</code></pre>\n<pre><code>Reducing memory usage for: train\n--- Memory usage before: 3374.26 MB\n--- Memory usage after: 843.57 MB\n--- Decreased memory usage by 75.0%\n\nReducing memory usage for: test\n--- Memory usage before: 3448.84 MB\n--- Memory usage after: 862.21 MB\n--- Decreased memory usage by 75.0%\n</code></pre>",
  "messages": [
    {
      "id": "3208526",
      "postDate": "05/24/2025 09:42:57",
      "content": "<p>The dataset for this competition is very large, both in terms of rows and columns, and it's likely to run into memory issues during EDA or model training. There are many ways to handle such challenges, but one simple  method is to reduce the memory footprint of the dataset. This can be done by changing data types to lower-precision formats, which helps save memory space. Below is a code snippet that I frequently use when working with large datasets on Kaggle.</p>\n<p>Hope you find it useful, and best of luck with the competition!</p>\n<pre><code> ():    \n    (, dataset)\n    initial_mem_usage = dataframe.memory_usage().() / **\n\n     col  dataframe.columns:\n        col_type = dataframe[col].dtype\n\n        c_min = dataframe[col].()\n        c_max = dataframe[col].()\n         (col_type)[:] == :\n             c_min &gt; np.iinfo(np.int8).  c_max &lt; np.iinfo(np.int8).:\n                dataframe[col] = dataframe[col].astype(np.int8)\n             c_min &gt; np.iinfo(np.int16).  c_max &lt; np.iinfo(np.int16).:\n                dataframe[col] = dataframe[col].astype(np.int16)\n             c_min &gt; np.iinfo(np.int32).  c_max &lt; np.iinfo(np.int32).:\n                dataframe[col] = dataframe[col].astype(np.int32)\n             c_min &gt; np.iinfo(np.int64).  c_max &lt; np.iinfo(np.int64).:\n                dataframe[col] = dataframe[col].astype(np.int64)\n        :\n             c_min &gt; np.finfo(np.float16).  c_max &lt; np.finfo(np.float16).:\n                dataframe[col] = dataframe[col].astype(np.float16)\n             c_min &gt; np.finfo(np.float32).  c_max &lt; np.finfo(np.float32).:\n                dataframe[col] = dataframe[col].astype(np.float32)\n            :\n                dataframe[col] = dataframe[col].astype(np.float64)\n\n    final_mem_usage = dataframe.memory_usage().() / **\n    (.(initial_mem_usage))\n    (.(final_mem_usage))\n    (.( * (initial_mem_usage - final_mem_usage) / initial_mem_usage))\n\n     dataframe\n</code></pre>\n<pre><code>Reducing memory usage for: train\n--- Memory usage before: 3374.26 MB\n--- Memory usage after: 843.57 MB\n--- Decreased memory usage by 75.0%\n\nReducing memory usage for: test\n--- Memory usage before: 3448.84 MB\n--- Memory usage after: 862.21 MB\n--- Decreased memory usage by 75.0%\n</code></pre>",
      "rawMarkdown": "The dataset for this competition is very large, both in terms of rows and columns, and it's likely to run into memory issues during EDA or model training. There are many ways to handle such challenges, but one simple  method is to reduce the memory footprint of the dataset. This can be done by changing data types to lower-precision formats, which helps save memory space. Below is a code snippet that I frequently use when working with large datasets on Kaggle.\n\nHope you find it useful, and best of luck with the competition!\n\n```python\ndef reduce_mem_usage(dataframe, dataset):    \n    print('Reducing memory usage for:', dataset)\n    initial_mem_usage = dataframe.memory_usage().sum() / 1024**2\n    \n    for col in dataframe.columns:\n        col_type = dataframe[col].dtype\n\n        c_min = dataframe[col].min()\n        c_max = dataframe[col].max()\n        if str(col_type)[:3] == 'int':\n            if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                dataframe[col] = dataframe[col].astype(np.int8)\n            elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                dataframe[col] = dataframe[col].astype(np.int16)\n            elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                dataframe[col] = dataframe[col].astype(np.int32)\n            elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                dataframe[col] = dataframe[col].astype(np.int64)\n        else:\n            if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                dataframe[col] = dataframe[col].astype(np.float16)\n            elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                dataframe[col] = dataframe[col].astype(np.float32)\n            else:\n                dataframe[col] = dataframe[col].astype(np.float64)\n\n    final_mem_usage = dataframe.memory_usage().sum() / 1024**2\n    print('--- Memory usage before: {:.2f} MB'.format(initial_mem_usage))\n    print('--- Memory usage after: {:.2f} MB'.format(final_mem_usage))\n    print('--- Decreased memory usage by {:.1f}%\\n'.format(100 * (initial_mem_usage - final_mem_usage) / initial_mem_usage))\n\n    return dataframe\n```\n```text\nReducing memory usage for: train\n--- Memory usage before: 3374.26 MB\n--- Memory usage after: 843.57 MB\n--- Decreased memory usage by 75.0%\n\nReducing memory usage for: test\n--- Memory usage before: 3448.84 MB\n--- Memory usage after: 862.21 MB\n--- Decreased memory usage by 75.0%\n```",
      "votes": null
    },
    {
      "id": "3208559",
      "postDate": "05/24/2025 10:58:56",
      "content": "<p><a href=\"https://www.kaggle.com/ravaghi\" target=\"_blank\">@ravaghi</a> one can do this in one line as below-</p>\n<pre><code> polars  pl\ndf.select(pl.().shrink_dtype())\n</code></pre>",
      "rawMarkdown": "ravaghi one can do this in one line as below-\n```python\nimport polars as pl\ndf.select(pl.all().shrink_dtype())\n```",
      "votes": null
    },
    {
      "id": "3208594",
      "postDate": "05/24/2025 11:44:48",
      "content": "<p>Thanks for the tip! I might be doing something wrong since I don't use Polars as much as Pandas, but with the following code, I get only a 50% reduction in memory usage compared to the 75% reduction I get with Pandas.</p>\n<pre><code> ():    \n    (, dataset)\n    initial_mem_usage = dataframe.estimated_size() / **\n\n    dataframe = dataframe.select(pl.().shrink_dtype())\n\n    final_mem_usage = dataframe.estimated_size() / **\n    (.(initial_mem_usage))\n    (.(final_mem_usage))\n    (.( * (initial_mem_usage - final_mem_usage) / initial_mem_usage))\n\n     dataframe\n</code></pre>\n<pre><code>Reducing memory usage for: train\n--- Memory usage before: 3598.94 MB\n--- Memory usage after: 1801.48 MB\n--- Decreased memory usage by 49.9%\n</code></pre>",
      "rawMarkdown": "Thanks for the tip! I might be doing something wrong since I don't use Polars as much as Pandas, but with the following code, I get only a 50% reduction in memory usage compared to the 75% reduction I get with Pandas.\n```python\ndef reduce_mem_usage(dataframe, dataset):    \n    print('Reducing memory usage for:', dataset)\n    initial_mem_usage = dataframe.estimated_size() / 1024**2\n    \n    dataframe = dataframe.select(pl.all().shrink_dtype())\n\n    final_mem_usage = dataframe.estimated_size() / 1024**2\n    print('--- Memory usage before: {:.2f} MB'.format(initial_mem_usage))\n    print('--- Memory usage after: {:.2f} MB'.format(final_mem_usage))\n    print('--- Decreased memory usage by {:.1f}%\\n'.format(100 * (initial_mem_usage - final_mem_usage) / initial_mem_usage))\n\n    return dataframe\n```\n```text\nReducing memory usage for: train\n--- Memory usage before: 3598.94 MB\n--- Memory usage after: 1801.48 MB\n--- Decreased memory usage by 49.9%\n```",
      "votes": null
    },
    {
      "id": "3208775",
      "postDate": "05/24/2025 16:56:35",
      "content": "<p>Do you have an experiment to test how feature‘s precision impacts the model's performance?</p>",
      "rawMarkdown": "Do you have an experiment to test how feature‘s precision impacts the model's performance?",
      "votes": null
    },
    {
      "id": "3208964",
      "postDate": "05/25/2025 03:11:01",
      "content": "<p>my solution is using batchs to load and process data (pretty bad one though)<br>\non train:</p>\n<pre><code>batch_num = \n i  (batch_num):\n    new_df = pd.DataFrame()\n    ()\n    n = (train)\n    train_batch = train[(i / batch_num * n): ((i + ) / batch_num * n)]\n    train_batch = features(train_batch)\n    new_df = pd.concat([new_df, train_batch], axis = )\n\ntrain = new_df`\n</code></pre>\n<p>on test</p>\n<pre><code> ():\n    results = []\n     i  (batch_num):\n        ()\n        n = (test)\n        test_batch = test[((i) / batch_num * n): ((i + ) / batch_num * n)]\n        test_batch = features(test_batch, pca)\n        pred_batch = evaluate(test_batch, lightgbm_model)\n        results.extend(pred_batch)\n     results\n</code></pre>",
      "rawMarkdown": "my solution is using batchs to load and process data (pretty bad one though)\non train:\n```python\nbatch_num = 10\nfor i in range(batch_num):\n    new_df = pd.DataFrame()\n    print(f\"Add feature for batch {i+1} / {batch_num}\")\n    n = len(train)\n    train_batch = train[int(i / batch_num * n): int((i + 1) / batch_num * n)]\n    train_batch = features(train_batch)\n    new_df = pd.concat([new_df, train_batch], axis = 0)\n\ntrain = new_df`\n```\non test\n```python\ndef predict(test, batch_num = 10):\n    results = []\n    for i in range(batch_num):\n        print(f'Evaluating batch {i+1} / {batch_num}')\n        n = len(test)\n        test_batch = test[int((i) / batch_num * n): int((i + 1) / batch_num * n)]\n        test_batch = features(test_batch, pca)\n        pred_batch = evaluate(test_batch, lightgbm_model)\n        results.extend(pred_batch)\n    return results\n```",
      "votes": null
    },
    {
      "id": "3209014",
      "postDate": "05/25/2025 05:26:20",
      "content": "<p><a href=\"https://www.kaggle.com/ivantang86\" target=\"_blank\">@ivantang86</a> you can save the train and test data in polars using parquet and <strong>hive storage</strong> using partition id. Retrieval is also easy-</p>\n<pre><code> polars  pl\ndf.write_parquet(, partition_id = )\n</code></pre>\n<pre><code> polars  pl\npl.scan_parquet(df) -- retrieves the whole file  lazy mode\npl.read_parquet(df) -- retrieves the whole file  eager mode\npl.scan_parquet(df, hive_partitioning = , n_rows = , row_index_offset=) -- examples of selective lazy imports  hive files\npl.scan_parquet(df).(pl.col() &lt;= cutoff) -- another way to read  the file  lazy mode\n</code></pre>",
      "rawMarkdown": "ivantang86 you can save the train and test data in polars using parquet and **hive storage** using partition id. Retrieval is also easy-\n\n```python\nimport polars as pl\ndf.write_parquet(\"myparquet.parquet\", partition_id = \"mypartitionid_column\")\n```\n\n```python\nimport polars as pl\npl.scan_parquet(df) -- retrieves the whole file in lazy mode\npl.read_parquet(df) -- retrieves the whole file in eager mode\npl.scan_parquet(df, hive_partitioning = True, n_rows = , row_index_offset=) -- examples of selective lazy imports from hive files\npl.scan_parquet(df).filter(pl.col(\"mypartitionid_column\") <= cutoff) -- another way to read in the file with lazy mode\n```",
      "votes": null
    },
    {
      "id": "3209072",
      "postDate": "05/25/2025 07:27:46",
      "content": "<p><a href=\"https://www.kaggle.com/andrewguanzc\" target=\"_blank\">@andrewguanzc</a> I don't have a powerful computer, so I haven't tested with full precision. Besides, lower precisions are giving decent CV and LB scores, so I'm not worried about losing performance.</p>",
      "rawMarkdown": "andrewguanzc I don't have a powerful computer, so I haven't tested with full precision. Besides, lower precisions are giving decent CV and LB scores, so I'm not worried about losing performance.",
      "votes": null
    },
    {
      "id": "3209117",
      "postDate": "05/25/2025 08:56:11",
      "content": "<p>Thanks for the post.</p>",
      "rawMarkdown": "Thanks for the post.",
      "votes": null
    },
    {
      "id": "3209374",
      "postDate": "05/25/2025 17:08:57",
      "content": "<p>Changing the data type of the feature is the method i use most and it work fine. Reducing the memory usage. <br>\nLike if there is a column of \"age\" then by default it is int64 which is unnecessary, so i change it to int8 or int16 depending on the situation.</p>",
      "rawMarkdown": "Changing the data type of the feature is the method i use most and it work fine. Reducing the memory usage. \nLike if there is a column of \"age\" then by default it is int64 which is unnecessary, so i change it to int8 or int16 depending on the situation.",
      "votes": null
    },
    {
      "id": "3209377",
      "postDate": "05/25/2025 17:12:55",
      "content": "<p>Sometimes I just convert the datatype by my own without using any library but sometimes I use polars library.</p>",
      "rawMarkdown": "Sometimes I just convert the datatype by my own without using any library but sometimes I use polars library.",
      "votes": null
    },
    {
      "id": "3209402",
      "postDate": "05/25/2025 18:12:50",
      "content": "<p>Excellent memory optimization guide <a href=\"https://www.kaggle.com/ravaghi\" target=\"_blank\">@ravaghi</a>! 🚀 Building on your fantastic foundation, here are some additional advanced techniques for handling large-scale crypto datasets:</p>\n<p><strong>🔧 Advanced Memory Optimization Strategies:</strong></p>\n<p><strong>1. Chunked Processing with Dask:</strong></p>\n<pre><code> dask.dataframe  dd\n\n\ndf = dd.read_csv(, blocksize=)\nresult = df.groupby().agg({: }).compute()\n</code></pre>\n<p><strong>2. Memory-Mapped Files for Ultra-Large Datasets:</strong></p>\n<pre><code> numpy  np\n numpy.lib.  open_memmap\n\n\nfeatures_mmap = open_memmap(, dtype=, mode=, shape=(n_samples, n_features))\n\n</code></pre>\n<p><strong>3. Feature Store Approach:</strong></p>\n<pre><code>\n h5py\n\n h5py.File(, )  f:\n    f.create_dataset(, data=train_features, compression=)\n    f.create_dataset(, data=train_targets, compression=)\n</code></pre>\n<p><strong>4. Smart Categorical Encoding:</strong></p>\n<pre><code>\ndf[] = df[].astype()\ndf[] = df[].astype()\n\n</code></pre>\n<p><strong>5. Incremental Learning Pipeline:</strong></p>\n<pre><code> sklearn.linear_model  SGDRegressor\n\n\nmodel = SGDRegressor()\n chunk  pd.read_csv(, chunksize=):\n    chunk = preprocess(chunk)\n    model.partial_fit(chunk[features], chunk[target])\n</code></pre>\n<p><strong>6. GPU Memory Management (if available):</strong></p>\n<pre><code> cudf  \n\n\ndf_gpu = cudf.read_csv()\ndf_gpu = df_gpu.astype({: , : })\n</code></pre>\n<p><strong>💡 Pro Tips for Crypto Data:</strong></p>\n<ul>\n<li><strong>Time-based chunking</strong>: Process data by date ranges to maintain temporal relationships</li>\n<li><strong>Symbol-based partitioning</strong>: Store each crypto symbol separately for parallel processing</li>\n<li><strong>Feature caching</strong>: Cache expensive feature computations using joblib.Memory</li>\n<li><strong>Lazy evaluation</strong>: Use Polars' lazy API for query optimization before execution</li>\n</ul>\n<p><strong>Memory Monitoring:</strong></p>\n<pre><code> psutil\n gc\n\n ():\n    process = psutil.Process()\n    ()\n\n\ngc.collect()\n</code></pre>\n<p>The combination of your dtype optimization + these techniques can handle datasets 10x larger than available RAM! 📊</p>\n<p>For this crypto competition specifically, I'd recommend:</p>\n<ol>\n<li>Your dtype reduction (75% savings)</li>\n<li>Time-based chunking for feature engineering</li>\n<li>Parquet storage with compression</li>\n<li>Incremental model training</li>\n</ol>\n<p>Great foundation post - these memory challenges are where competitions are won! 💪</p>",
      "rawMarkdown": "Excellent memory optimization guide @ravaghi! 🚀 Building on your fantastic foundation, here are some additional advanced techniques for handling large-scale crypto datasets:\n\n**🔧 Advanced Memory Optimization Strategies:**\n\n**1. Chunked Processing with Dask:**\n```python\nimport dask.dataframe as dd\n\n# Process data in chunks without loading everything into memory\ndf = dd.read_csv('large_crypto_data.csv', blocksize='100MB')\nresult = df.groupby('symbol').agg({'price': 'mean'}).compute()\n```\n\n**2. Memory-Mapped Files for Ultra-Large Datasets:**\n```python\nimport numpy as np\nfrom numpy.lib.format import open_memmap\n\n# Create memory-mapped array for features\nfeatures_mmap = open_memmap('features.dat', dtype='float32', mode='w+', shape=(n_samples, n_features))\n# Access like regular array but stays on disk\n```\n\n**3. Feature Store Approach:**\n```python\n# Store preprocessed features in HDF5 for fast random access\nimport h5py\n\nwith h5py.File('crypto_features.h5', 'w') as f:\n    f.create_dataset('train_features', data=train_features, compression='gzip')\n    f.create_dataset('train_targets', data=train_targets, compression='gzip')\n```\n\n**4. Smart Categorical Encoding:**\n```python\n# Use category dtype for string columns (massive memory savings)\ndf['symbol'] = df['symbol'].astype('category')\ndf['exchange'] = df['exchange'].astype('category')\n# Can reduce memory by 80%+ for high-cardinality categoricals\n```\n\n**5. Incremental Learning Pipeline:**\n```python\nfrom sklearn.linear_model import SGDRegressor\n\n# Train models incrementally on data chunks\nmodel = SGDRegressor()\nfor chunk in pd.read_csv('data.csv', chunksize=10000):\n    chunk = preprocess(chunk)\n    model.partial_fit(chunk[features], chunk[target])\n```\n\n**6. GPU Memory Management (if available):**\n```python\nimport cudf  # RAPIDS for GPU acceleration\n\n# Process on GPU with automatic memory management\ndf_gpu = cudf.read_csv('data.csv')\ndf_gpu = df_gpu.astype({'price': 'float32', 'volume': 'float32'})\n```\n\n**💡 Pro Tips for Crypto Data:**\n- **Time-based chunking**: Process data by date ranges to maintain temporal relationships\n- **Symbol-based partitioning**: Store each crypto symbol separately for parallel processing\n- **Feature caching**: Cache expensive feature computations using joblib.Memory\n- **Lazy evaluation**: Use Polars' lazy API for query optimization before execution\n\n**Memory Monitoring:**\n```python\nimport psutil\nimport gc\n\ndef monitor_memory():\n    process = psutil.Process()\n    print(f\"Memory usage: {process.memory_info().rss / 1024 / 1024:.2f} MB\")\n    \n# Force garbage collection after processing chunks\ngc.collect()\n```\n\nThe combination of your dtype optimization + these techniques can handle datasets 10x larger than available RAM! 📊\n\nFor this crypto competition specifically, I'd recommend:\n1. Your dtype reduction (75% savings)\n2. Time-based chunking for feature engineering\n3. Parquet storage with compression\n4. Incremental model training\n\nGreat foundation post - these memory challenges are where competitions are won! 💪",
      "votes": null
    },
    {
      "id": "3209420",
      "postDate": "05/25/2025 19:01:18",
      "content": "<p>Thank you, i learnt much from the techniques you gave </p>",
      "rawMarkdown": "Thank you, i learnt much from the techniques you gave",
      "votes": null
    },
    {
      "id": "3211335",
      "postDate": "05/28/2025 10:11:36",
      "content": "<p>This is quite a handy tip! Thanks for sharing it.</p>",
      "rawMarkdown": "This is quite a handy tip! Thanks for sharing it.",
      "votes": null
    },
    {
      "id": "3211709",
      "postDate": "05/28/2025 18:27:38",
      "content": "<p>Do you have an experiment to test how feature‘s precision impacts the overall model's performance?</p>",
      "rawMarkdown": "Do you have an experiment to test how feature‘s precision impacts the overall model's performance?",
      "votes": null
    },
    {
      "id": "3231091",
      "postDate": "06/23/2025 22:07:48",
      "content": "<p>I think polar doesn't work out so well here, because later you have to convert it back to pandas, which will undo most of the memory saving.</p>",
      "rawMarkdown": "I think polar doesn't work out so well here, because later you have to convert it back to pandas, which will undo most of the memory saving.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3208559,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "05/24/2025 10:58:56",
      "content": "<p><a href=\"https://www.kaggle.com/ravaghi\" target=\"_blank\">@ravaghi</a> one can do this in one line as below-</p>\n<pre><code> polars  pl\ndf.select(pl.().shrink_dtype())\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3208594,
          "author_name": "ravaghi",
          "author_url": "",
          "post_date": "05/24/2025 11:44:48",
          "content": "<p>Thanks for the tip! I might be doing something wrong since I don't use Polars as much as Pandas, but with the following code, I get only a 50% reduction in memory usage compared to the 75% reduction I get with Pandas.</p>\n<pre><code> ():    \n    (, dataset)\n    initial_mem_usage = dataframe.estimated_size() / **\n\n    dataframe = dataframe.select(pl.().shrink_dtype())\n\n    final_mem_usage = dataframe.estimated_size() / **\n    (.(initial_mem_usage))\n    (.(final_mem_usage))\n    (.( * (initial_mem_usage - final_mem_usage) / initial_mem_usage))\n\n     dataframe\n</code></pre>\n<pre><code>Reducing memory usage for: train\n--- Memory usage before: 3598.94 MB\n--- Memory usage after: 1801.48 MB\n--- Decreased memory usage by 49.9%\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 3208775,
              "author_name": "andrewguanzc",
              "author_url": "",
              "post_date": "05/24/2025 16:56:35",
              "content": "<p>Do you have an experiment to test how feature‘s precision impacts the model's performance?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3209072,
                  "author_name": "ravaghi",
                  "author_url": "",
                  "post_date": "05/25/2025 07:27:46",
                  "content": "<p><a href=\"https://www.kaggle.com/andrewguanzc\" target=\"_blank\">@andrewguanzc</a> I don't have a powerful computer, so I haven't tested with full precision. Besides, lower precisions are giving decent CV and LB scores, so I'm not worried about losing performance.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 3231091,
              "author_name": "andywangww",
              "author_url": "",
              "post_date": "06/23/2025 22:07:48",
              "content": "<p>I think polar doesn't work out so well here, because later you have to convert it back to pandas, which will undo most of the memory saving.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3211335,
          "author_name": "dalloliogm",
          "author_url": "",
          "post_date": "05/28/2025 10:11:36",
          "content": "<p>This is quite a handy tip! Thanks for sharing it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3208964,
      "author_name": "ivantang86",
      "author_url": "",
      "post_date": "05/25/2025 03:11:01",
      "content": "<p>my solution is using batchs to load and process data (pretty bad one though)<br>\non train:</p>\n<pre><code>batch_num = \n i  (batch_num):\n    new_df = pd.DataFrame()\n    ()\n    n = (train)\n    train_batch = train[(i / batch_num * n): ((i + ) / batch_num * n)]\n    train_batch = features(train_batch)\n    new_df = pd.concat([new_df, train_batch], axis = )\n\ntrain = new_df`\n</code></pre>\n<p>on test</p>\n<pre><code> ():\n    results = []\n     i  (batch_num):\n        ()\n        n = (test)\n        test_batch = test[((i) / batch_num * n): ((i + ) / batch_num * n)]\n        test_batch = features(test_batch, pca)\n        pred_batch = evaluate(test_batch, lightgbm_model)\n        results.extend(pred_batch)\n     results\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3209014,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "05/25/2025 05:26:20",
          "content": "<p><a href=\"https://www.kaggle.com/ivantang86\" target=\"_blank\">@ivantang86</a> you can save the train and test data in polars using parquet and <strong>hive storage</strong> using partition id. Retrieval is also easy-</p>\n<pre><code> polars  pl\ndf.write_parquet(, partition_id = )\n</code></pre>\n<pre><code> polars  pl\npl.scan_parquet(df) -- retrieves the whole file  lazy mode\npl.read_parquet(df) -- retrieves the whole file  eager mode\npl.scan_parquet(df, hive_partitioning = , n_rows = , row_index_offset=) -- examples of selective lazy imports  hive files\npl.scan_parquet(df).(pl.col() &lt;= cutoff) -- another way to read  the file  lazy mode\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3209117,
      "author_name": "ladiposamson",
      "author_url": "",
      "post_date": "05/25/2025 08:56:11",
      "content": "<p>Thanks for the post.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3209374,
      "author_name": "sharmajicoder",
      "author_url": "",
      "post_date": "05/25/2025 17:08:57",
      "content": "<p>Changing the data type of the feature is the method i use most and it work fine. Reducing the memory usage. <br>\nLike if there is a column of \"age\" then by default it is int64 which is unnecessary, so i change it to int8 or int16 depending on the situation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3209377,
      "author_name": "sharmajicoder",
      "author_url": "",
      "post_date": "05/25/2025 17:12:55",
      "content": "<p>Sometimes I just convert the datatype by my own without using any library but sometimes I use polars library.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3209402,
      "author_name": "harshithvaddiparthy",
      "author_url": "",
      "post_date": "05/25/2025 18:12:50",
      "content": "<p>Excellent memory optimization guide <a href=\"https://www.kaggle.com/ravaghi\" target=\"_blank\">@ravaghi</a>! 🚀 Building on your fantastic foundation, here are some additional advanced techniques for handling large-scale crypto datasets:</p>\n<p><strong>🔧 Advanced Memory Optimization Strategies:</strong></p>\n<p><strong>1. Chunked Processing with Dask:</strong></p>\n<pre><code> dask.dataframe  dd\n\n\ndf = dd.read_csv(, blocksize=)\nresult = df.groupby().agg({: }).compute()\n</code></pre>\n<p><strong>2. Memory-Mapped Files for Ultra-Large Datasets:</strong></p>\n<pre><code> numpy  np\n numpy.lib.  open_memmap\n\n\nfeatures_mmap = open_memmap(, dtype=, mode=, shape=(n_samples, n_features))\n\n</code></pre>\n<p><strong>3. Feature Store Approach:</strong></p>\n<pre><code>\n h5py\n\n h5py.File(, )  f:\n    f.create_dataset(, data=train_features, compression=)\n    f.create_dataset(, data=train_targets, compression=)\n</code></pre>\n<p><strong>4. Smart Categorical Encoding:</strong></p>\n<pre><code>\ndf[] = df[].astype()\ndf[] = df[].astype()\n\n</code></pre>\n<p><strong>5. Incremental Learning Pipeline:</strong></p>\n<pre><code> sklearn.linear_model  SGDRegressor\n\n\nmodel = SGDRegressor()\n chunk  pd.read_csv(, chunksize=):\n    chunk = preprocess(chunk)\n    model.partial_fit(chunk[features], chunk[target])\n</code></pre>\n<p><strong>6. GPU Memory Management (if available):</strong></p>\n<pre><code> cudf  \n\n\ndf_gpu = cudf.read_csv()\ndf_gpu = df_gpu.astype({: , : })\n</code></pre>\n<p><strong>💡 Pro Tips for Crypto Data:</strong></p>\n<ul>\n<li><strong>Time-based chunking</strong>: Process data by date ranges to maintain temporal relationships</li>\n<li><strong>Symbol-based partitioning</strong>: Store each crypto symbol separately for parallel processing</li>\n<li><strong>Feature caching</strong>: Cache expensive feature computations using joblib.Memory</li>\n<li><strong>Lazy evaluation</strong>: Use Polars' lazy API for query optimization before execution</li>\n</ul>\n<p><strong>Memory Monitoring:</strong></p>\n<pre><code> psutil\n gc\n\n ():\n    process = psutil.Process()\n    ()\n\n\ngc.collect()\n</code></pre>\n<p>The combination of your dtype optimization + these techniques can handle datasets 10x larger than available RAM! 📊</p>\n<p>For this crypto competition specifically, I'd recommend:</p>\n<ol>\n<li>Your dtype reduction (75% savings)</li>\n<li>Time-based chunking for feature engineering</li>\n<li>Parquet storage with compression</li>\n<li>Incremental model training</li>\n</ol>\n<p>Great foundation post - these memory challenges are where competitions are won! 💪</p>",
      "votes": null,
      "replies": [
        {
          "id": 3209420,
          "author_name": "chikonzeroselemani",
          "author_url": "",
          "post_date": "05/25/2025 19:01:18",
          "content": "<p>Thank you, i learnt much from the techniques you gave </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3211709,
      "author_name": "icebearogo",
      "author_url": "",
      "post_date": "05/28/2025 18:27:38",
      "content": "<p>Do you have an experiment to test how feature‘s precision impacts the overall model's performance?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3208526": "The dataset for this competition is very large, both in terms of rows and columns, and it's likely to run into memory issues during EDA or model training. There are many ways to handle such challenges, but one simple  method is to reduce the memory footprint of the dataset. This can be done by changing data types to lower-precision formats, which helps save memory space. Below is a code snippet that I frequently use when working with large datasets on Kaggle.\n\nHope you find it useful, and best of luck with the competition!\n\n```python\ndef reduce_mem_usage(dataframe, dataset):    \n    print('Reducing memory usage for:', dataset)\n    initial_mem_usage = dataframe.memory_usage().sum() / 1024**2\n    \n    for col in dataframe.columns:\n        col_type = dataframe[col].dtype\n\n        c_min = dataframe[col].min()\n        c_max = dataframe[col].max()\n        if str(col_type)[:3] == 'int':\n            if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                dataframe[col] = dataframe[col].astype(np.int8)\n            elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                dataframe[col] = dataframe[col].astype(np.int16)\n            elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                dataframe[col] = dataframe[col].astype(np.int32)\n            elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                dataframe[col] = dataframe[col].astype(np.int64)\n        else:\n            if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                dataframe[col] = dataframe[col].astype(np.float16)\n            elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                dataframe[col] = dataframe[col].astype(np.float32)\n            else:\n                dataframe[col] = dataframe[col].astype(np.float64)\n\n    final_mem_usage = dataframe.memory_usage().sum() / 1024**2\n    print('--- Memory usage before: {:.2f} MB'.format(initial_mem_usage))\n    print('--- Memory usage after: {:.2f} MB'.format(final_mem_usage))\n    print('--- Decreased memory usage by {:.1f}%\\n'.format(100 * (initial_mem_usage - final_mem_usage) / initial_mem_usage))\n\n    return dataframe\n```\n```text\nReducing memory usage for: train\n--- Memory usage before: 3374.26 MB\n--- Memory usage after: 843.57 MB\n--- Decreased memory usage by 75.0%\n\nReducing memory usage for: test\n--- Memory usage before: 3448.84 MB\n--- Memory usage after: 862.21 MB\n--- Decreased memory usage by 75.0%\n```",
    "3208559": "ravaghi one can do this in one line as below-\n```python\nimport polars as pl\ndf.select(pl.all().shrink_dtype())\n```",
    "3208594": "Thanks for the tip! I might be doing something wrong since I don't use Polars as much as Pandas, but with the following code, I get only a 50% reduction in memory usage compared to the 75% reduction I get with Pandas.\n```python\ndef reduce_mem_usage(dataframe, dataset):    \n    print('Reducing memory usage for:', dataset)\n    initial_mem_usage = dataframe.estimated_size() / 1024**2\n    \n    dataframe = dataframe.select(pl.all().shrink_dtype())\n\n    final_mem_usage = dataframe.estimated_size() / 1024**2\n    print('--- Memory usage before: {:.2f} MB'.format(initial_mem_usage))\n    print('--- Memory usage after: {:.2f} MB'.format(final_mem_usage))\n    print('--- Decreased memory usage by {:.1f}%\\n'.format(100 * (initial_mem_usage - final_mem_usage) / initial_mem_usage))\n\n    return dataframe\n```\n```text\nReducing memory usage for: train\n--- Memory usage before: 3598.94 MB\n--- Memory usage after: 1801.48 MB\n--- Decreased memory usage by 49.9%\n```",
    "3208775": "Do you have an experiment to test how feature‘s precision impacts the model's performance?",
    "3208964": "my solution is using batchs to load and process data (pretty bad one though)\non train:\n```python\nbatch_num = 10\nfor i in range(batch_num):\n    new_df = pd.DataFrame()\n    print(f\"Add feature for batch {i+1} / {batch_num}\")\n    n = len(train)\n    train_batch = train[int(i / batch_num * n): int((i + 1) / batch_num * n)]\n    train_batch = features(train_batch)\n    new_df = pd.concat([new_df, train_batch], axis = 0)\n\ntrain = new_df`\n```\non test\n```python\ndef predict(test, batch_num = 10):\n    results = []\n    for i in range(batch_num):\n        print(f'Evaluating batch {i+1} / {batch_num}')\n        n = len(test)\n        test_batch = test[int((i) / batch_num * n): int((i + 1) / batch_num * n)]\n        test_batch = features(test_batch, pca)\n        pred_batch = evaluate(test_batch, lightgbm_model)\n        results.extend(pred_batch)\n    return results\n```",
    "3209014": "ivantang86 you can save the train and test data in polars using parquet and **hive storage** using partition id. Retrieval is also easy-\n\n```python\nimport polars as pl\ndf.write_parquet(\"myparquet.parquet\", partition_id = \"mypartitionid_column\")\n```\n\n```python\nimport polars as pl\npl.scan_parquet(df) -- retrieves the whole file in lazy mode\npl.read_parquet(df) -- retrieves the whole file in eager mode\npl.scan_parquet(df, hive_partitioning = True, n_rows = , row_index_offset=) -- examples of selective lazy imports from hive files\npl.scan_parquet(df).filter(pl.col(\"mypartitionid_column\") <= cutoff) -- another way to read in the file with lazy mode\n```",
    "3209072": "andrewguanzc I don't have a powerful computer, so I haven't tested with full precision. Besides, lower precisions are giving decent CV and LB scores, so I'm not worried about losing performance.",
    "3209117": "Thanks for the post.",
    "3209374": "Changing the data type of the feature is the method i use most and it work fine. Reducing the memory usage. \nLike if there is a column of \"age\" then by default it is int64 which is unnecessary, so i change it to int8 or int16 depending on the situation.",
    "3209377": "Sometimes I just convert the datatype by my own without using any library but sometimes I use polars library.",
    "3209402": "Excellent memory optimization guide @ravaghi! 🚀 Building on your fantastic foundation, here are some additional advanced techniques for handling large-scale crypto datasets:\n\n**🔧 Advanced Memory Optimization Strategies:**\n\n**1. Chunked Processing with Dask:**\n```python\nimport dask.dataframe as dd\n\n# Process data in chunks without loading everything into memory\ndf = dd.read_csv('large_crypto_data.csv', blocksize='100MB')\nresult = df.groupby('symbol').agg({'price': 'mean'}).compute()\n```\n\n**2. Memory-Mapped Files for Ultra-Large Datasets:**\n```python\nimport numpy as np\nfrom numpy.lib.format import open_memmap\n\n# Create memory-mapped array for features\nfeatures_mmap = open_memmap('features.dat', dtype='float32', mode='w+', shape=(n_samples, n_features))\n# Access like regular array but stays on disk\n```\n\n**3. Feature Store Approach:**\n```python\n# Store preprocessed features in HDF5 for fast random access\nimport h5py\n\nwith h5py.File('crypto_features.h5', 'w') as f:\n    f.create_dataset('train_features', data=train_features, compression='gzip')\n    f.create_dataset('train_targets', data=train_targets, compression='gzip')\n```\n\n**4. Smart Categorical Encoding:**\n```python\n# Use category dtype for string columns (massive memory savings)\ndf['symbol'] = df['symbol'].astype('category')\ndf['exchange'] = df['exchange'].astype('category')\n# Can reduce memory by 80%+ for high-cardinality categoricals\n```\n\n**5. Incremental Learning Pipeline:**\n```python\nfrom sklearn.linear_model import SGDRegressor\n\n# Train models incrementally on data chunks\nmodel = SGDRegressor()\nfor chunk in pd.read_csv('data.csv', chunksize=10000):\n    chunk = preprocess(chunk)\n    model.partial_fit(chunk[features], chunk[target])\n```\n\n**6. GPU Memory Management (if available):**\n```python\nimport cudf  # RAPIDS for GPU acceleration\n\n# Process on GPU with automatic memory management\ndf_gpu = cudf.read_csv('data.csv')\ndf_gpu = df_gpu.astype({'price': 'float32', 'volume': 'float32'})\n```\n\n**💡 Pro Tips for Crypto Data:**\n- **Time-based chunking**: Process data by date ranges to maintain temporal relationships\n- **Symbol-based partitioning**: Store each crypto symbol separately for parallel processing\n- **Feature caching**: Cache expensive feature computations using joblib.Memory\n- **Lazy evaluation**: Use Polars' lazy API for query optimization before execution\n\n**Memory Monitoring:**\n```python\nimport psutil\nimport gc\n\ndef monitor_memory():\n    process = psutil.Process()\n    print(f\"Memory usage: {process.memory_info().rss / 1024 / 1024:.2f} MB\")\n    \n# Force garbage collection after processing chunks\ngc.collect()\n```\n\nThe combination of your dtype optimization + these techniques can handle datasets 10x larger than available RAM! 📊\n\nFor this crypto competition specifically, I'd recommend:\n1. Your dtype reduction (75% savings)\n2. Time-based chunking for feature engineering\n3. Parquet storage with compression\n4. Incremental model training\n\nGreat foundation post - these memory challenges are where competitions are won! 💪",
    "3209420": "Thank you, i learnt much from the techniques you gave",
    "3211335": "This is quite a handy tip! Thanks for sharing it.",
    "3211709": "Do you have an experiment to test how feature‘s precision impacts the overall model's performance?",
    "3231091": "I think polar doesn't work out so well here, because later you have to convert it back to pandas, which will undo most of the memory saving."
  },
  "source": "meta"
}