{
  "id": 382286,
  "title": "Hash trick to save Memory when merging data (64GB for all)",
  "url": "/competitions/otto-recommender-system/discussion/382286",
  "author_name": "",
  "post_date": "2023-01-30T13:04:16.249648900Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I played around the data of this competition using machine with 64G RAM, as you know even after saving data into chunks, using polars to replace pandas … OOM stills happens when you try to merge two large dataframes (like merging recall strategies, merge user features ..).</p>\n<p>I found a simple trick could be quite useful when merging two dataframes by session_id, the magic is hash</p>\n<pre><code> ():\n     (hashlib.md5(x.encode()).hexdigest(), ) % num_buckets\n</code></pre>\n<p>When saving the data into chunks, split the chunks with hash, for example for those session ids been hashed into 16th bucket <code>myhash_func(some_session_id) = 16</code>, save them into a file named 'bkt_0016.parquet', and keep doing these for all the data that need to be merged by session_id later.</p>\n<p>A great benefit of this, is that when merging dataframe A and B by session_id, instead of loading the whole dataframe A and B into memory, you can do this</p>\n<pre><code> f  glob.glob(dir_to_data_A):\n    file_name = f.split()[-]\n    A = pd.read_parquet(f)\n    B = pd.read_parquet(os.path.join(dir_to_data_B, file_name))\n    df = pd.merge(A, B, on=)\n    ....\n</code></pre>\n<p>This saves a lot of memory and also faster because merge is more efficient between small dataframes. I hashed all the data into 50 buckets, it works quite well for me.</p>",
  "messages": [
    {
      "id": "2121633",
      "postDate": "01/30/2023 13:04:16",
      "content": "<p>I played around the data of this competition using machine with 64G RAM, as you know even after saving data into chunks, using polars to replace pandas … OOM stills happens when you try to merge two large dataframes (like merging recall strategies, merge user features ..).</p>\n<p>I found a simple trick could be quite useful when merging two dataframes by session_id, the magic is hash</p>\n<pre><code> ():\n     (hashlib.md5(x.encode()).hexdigest(), ) % num_buckets\n</code></pre>\n<p>When saving the data into chunks, split the chunks with hash, for example for those session ids been hashed into 16th bucket <code>myhash_func(some_session_id) = 16</code>, save them into a file named 'bkt_0016.parquet', and keep doing these for all the data that need to be merged by session_id later.</p>\n<p>A great benefit of this, is that when merging dataframe A and B by session_id, instead of loading the whole dataframe A and B into memory, you can do this</p>\n<pre><code> f  glob.glob(dir_to_data_A):\n    file_name = f.split()[-]\n    A = pd.read_parquet(f)\n    B = pd.read_parquet(os.path.join(dir_to_data_B, file_name))\n    df = pd.merge(A, B, on=)\n    ....\n</code></pre>\n<p>This saves a lot of memory and also faster because merge is more efficient between small dataframes. I hashed all the data into 50 buckets, it works quite well for me.</p>",
      "rawMarkdown": "I played around the data of this competition using machine with 64G RAM, as you know even after saving data into chunks, using polars to replace pandas ... OOM stills happens when you try to merge two large dataframes (like merging recall strategies, merge user features ..).\n\nI found a simple trick could be quite useful when merging two dataframes by session_id, the magic is hash\n\n```python\ndef myhash_func(x, num_buckets):\n    return int(hashlib.md5(x.encode()).hexdigest(), 16) % num_buckets\n```\n\nWhen saving the data into chunks, split the chunks with hash, for example for those session ids been hashed into 16th bucket `myhash_func(some_session_id) = 16`, save them into a file named 'bkt_0016.parquet', and keep doing these for all the data that need to be merged by session_id later.\n\nA great benefit of this, is that when merging dataframe A and B by session_id, instead of loading the whole dataframe A and B into memory, you can do this\n\n```python\nfor f in glob.glob(dir_to_data_A):\n    file_name = f.split('/')[-1]\n    A = pd.read_parquet(f)\n    B = pd.read_parquet(os.path.join(dir_to_data_B, file_name))\n    df = pd.merge(A, B, on='session_id')\n    ....\n```\n\nThis saves a lot of memory and also faster because merge is more efficient between small dataframes. I hashed all the data into 50 buckets, it works quite well for me.",
      "votes": null
    },
    {
      "id": "2122115",
      "postDate": "01/30/2023 17:53:27",
      "content": "<p>Thanks for sharing your trick on using hash to save memory when merging data. Splitting the data into chunks and using the hash function to name the files is a smart way to reduce the memory footprint during merging. The method of reading and merging smaller chunks of data is much more efficient and less prone to OOM errors. Your technique of using 50 buckets and hashing the session IDs is a good idea that worked well for you.</p>\n<p>Additionally, another solution to save memory when merging large datasets is to use \"on-the-fly\" merging techniques, such as dask or vaex, which allow you to work with large datasets without having to load the entire dataset into memory. These libraries use a lazy evaluation approach, so you can perform operations on large datasets without having to load all the data into memory at once. With dask and vaex, you can perform operations like groupby, join, and aggregations in a memory-efficient manner, which can be a good alternative to the hash trick.</p>\n<p>The community can definitely benefit from your useful tip and explore other solutions to tackle the challenge of merging large datasets.</p>",
      "rawMarkdown": "Thanks for sharing your trick on using hash to save memory when merging data. Splitting the data into chunks and using the hash function to name the files is a smart way to reduce the memory footprint during merging. The method of reading and merging smaller chunks of data is much more efficient and less prone to OOM errors. Your technique of using 50 buckets and hashing the session IDs is a good idea that worked well for you.\n\nAdditionally, another solution to save memory when merging large datasets is to use \"on-the-fly\" merging techniques, such as dask or vaex, which allow you to work with large datasets without having to load the entire dataset into memory. These libraries use a lazy evaluation approach, so you can perform operations on large datasets without having to load all the data into memory at once. With dask and vaex, you can perform operations like groupby, join, and aggregations in a memory-efficient manner, which can be a good alternative to the hash trick.\n\nThe community can definitely benefit from your useful tip and explore other solutions to tackle the challenge of merging large datasets.",
      "votes": null
    },
    {
      "id": "2123509",
      "postDate": "01/31/2023 14:58:45",
      "content": "<p>Thanks for sharing dask and vaex, haven't used these two before, definitely will try next time !</p>",
      "rawMarkdown": "Thanks for sharing dask and vaex, haven't used these two before, definitely will try next time !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2122115,
      "author_name": "tahamhaider",
      "author_url": "",
      "post_date": "01/30/2023 17:53:27",
      "content": "<p>Thanks for sharing your trick on using hash to save memory when merging data. Splitting the data into chunks and using the hash function to name the files is a smart way to reduce the memory footprint during merging. The method of reading and merging smaller chunks of data is much more efficient and less prone to OOM errors. Your technique of using 50 buckets and hashing the session IDs is a good idea that worked well for you.</p>\n<p>Additionally, another solution to save memory when merging large datasets is to use \"on-the-fly\" merging techniques, such as dask or vaex, which allow you to work with large datasets without having to load the entire dataset into memory. These libraries use a lazy evaluation approach, so you can perform operations on large datasets without having to load all the data into memory at once. With dask and vaex, you can perform operations like groupby, join, and aggregations in a memory-efficient manner, which can be a good alternative to the hash trick.</p>\n<p>The community can definitely benefit from your useful tip and explore other solutions to tackle the challenge of merging large datasets.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2123509,
          "author_name": "kellenxu",
          "author_url": "",
          "post_date": "01/31/2023 14:58:45",
          "content": "<p>Thanks for sharing dask and vaex, haven't used these two before, definitely will try next time !</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2121633": "I played around the data of this competition using machine with 64G RAM, as you know even after saving data into chunks, using polars to replace pandas ... OOM stills happens when you try to merge two large dataframes (like merging recall strategies, merge user features ..).\n\nI found a simple trick could be quite useful when merging two dataframes by session_id, the magic is hash\n\n```python\ndef myhash_func(x, num_buckets):\n    return int(hashlib.md5(x.encode()).hexdigest(), 16) % num_buckets\n```\n\nWhen saving the data into chunks, split the chunks with hash, for example for those session ids been hashed into 16th bucket `myhash_func(some_session_id) = 16`, save them into a file named 'bkt_0016.parquet', and keep doing these for all the data that need to be merged by session_id later.\n\nA great benefit of this, is that when merging dataframe A and B by session_id, instead of loading the whole dataframe A and B into memory, you can do this\n\n```python\nfor f in glob.glob(dir_to_data_A):\n    file_name = f.split('/')[-1]\n    A = pd.read_parquet(f)\n    B = pd.read_parquet(os.path.join(dir_to_data_B, file_name))\n    df = pd.merge(A, B, on='session_id')\n    ....\n```\n\nThis saves a lot of memory and also faster because merge is more efficient between small dataframes. I hashed all the data into 50 buckets, it works quite well for me.",
    "2122115": "Thanks for sharing your trick on using hash to save memory when merging data. Splitting the data into chunks and using the hash function to name the files is a smart way to reduce the memory footprint during merging. The method of reading and merging smaller chunks of data is much more efficient and less prone to OOM errors. Your technique of using 50 buckets and hashing the session IDs is a good idea that worked well for you.\n\nAdditionally, another solution to save memory when merging large datasets is to use \"on-the-fly\" merging techniques, such as dask or vaex, which allow you to work with large datasets without having to load the entire dataset into memory. These libraries use a lazy evaluation approach, so you can perform operations on large datasets without having to load all the data into memory at once. With dask and vaex, you can perform operations like groupby, join, and aggregations in a memory-efficient manner, which can be a good alternative to the hash trick.\n\nThe community can definitely benefit from your useful tip and explore other solutions to tackle the challenge of merging large datasets.",
    "2123509": "Thanks for sharing dask and vaex, haven't used these two before, definitely will try next time !"
  },
  "source": "meta"
}