{
  "id": 207783,
  "title": "dump dictionary",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207783",
  "author_name": "qiaqia",
  "post_date": "2020-12-31T09:00:22.905000",
  "votes": 2,
  "comment_count": 14,
  "views": 0,
  "content": "<p>What format is normally used when using dump to compress dictionaries?<br>\nI converted the 6G dictionary to zip format memory overflow…<br>\nanother fomat can I choose?</p>",
  "messages": [
    {
      "id": 1133488,
      "postDate": "2020-12-31T09:00:22.907Z",
      "content": "<p>What format is normally used when using dump to compress dictionaries?<br>\nI converted the 6G dictionary to zip format memory overflow…<br>\nanother fomat can I choose?</p>",
      "rawMarkdown": "What format is normally used when using dump to compress dictionaries?\nI converted the 6G dictionary to zip format memory overflow...\nanother fomat can I choose?",
      "votes": 2
    },
    {
      "id": 1133513,
      "postDate": "2020-12-31T09:24:47.087Z",
      "content": "<p>You can use pickle to dump a dictionary, like this: </p>\n<pre><code>with open('dict', 'wb') as handle:\n    pickle.dump(dict, handle, protocol=pickle.HIGHEST_PROTOCOL)\n</code></pre>\n<p>If memory overflow still persists, you can use <a href=\"https://pypi.org/project/hickle/\" target=\"_blank\">hickle</a> (HDF5 alternative to pickle). It worked for me when the dict is too large, but it takes an awfully a long time to dump the dictionary.</p>",
      "rawMarkdown": "You can use pickle to dump a dictionary, like this: \n\n    with open('dict', 'wb') as handle:\n        pickle.dump(dict, handle, protocol=pickle.HIGHEST_PROTOCOL)\n\nIf memory overflow still persists, you can use [hickle](https://pypi.org/project/hickle/) (HDF5 alternative to pickle). It worked for me when the dict is too large, but it takes an awfully a long time to dump the dictionary.",
      "votes": 1,
      "replies": [
        {
          "id": 1133529,
          "postDate": "2020-12-31T09:36:47.627Z",
          "content": "<p>thank you,attempt dict is too large..</p>",
          "rawMarkdown": "thank you,attempt dict is too large.."
        }
      ]
    },
    {
      "id": 1133693,
      "postDate": "2020-12-31T12:46:23.937Z",
      "content": "<p>I use SQLite for nested dictionaries and pickle.dump /pickle.load for others. Pretty fast and it take only 1gb ram, when I load prepared features to kernel</p>",
      "rawMarkdown": "I use SQLite for nested dictionaries and pickle.dump /pickle.load for others. Pretty fast and it take only 1gb ram, when I load prepared features to kernel",
      "replies": [
        {
          "id": 1133697,
          "postDate": "2020-12-31T12:49:31.607Z",
          "content": "<p>P.S. Nested dictionaries take up most of the memory</p>",
          "rawMarkdown": "P.S. Nested dictionaries take up most of the memory"
        },
        {
          "id": 1133713,
          "postDate": "2020-12-31T13:10:02.903Z",
          "content": "<p>yes,nested dict use too much memory,Is sqlite fast to learn?I have never used this</p>",
          "rawMarkdown": "yes,nested dict use too much memory,Is sqlite fast to learn?I have never used this"
        },
        {
          "id": 1133714,
          "postDate": "2020-12-31T13:11:36.003Z",
          "content": "<p>Double nested dictionaries are just too big, normal nested dictionaries don't consume too much memory</p>",
          "rawMarkdown": "Double nested dictionaries are just too big, normal nested dictionaries don't consume too much memory"
        },
        {
          "id": 1133776,
          "postDate": "2020-12-31T14:18:08.050Z",
          "content": "<p>It looks like sql is the optimal solution for me, and I'm already consuming 11 G of memory in my inference notebook.. I can't emsemble saint model…</p>",
          "rawMarkdown": "It looks like sql is the optimal solution for me, and I'm already consuming 11 G of memory in my inference notebook.. I can't emsemble saint model..."
        },
        {
          "id": 1133798,
          "postDate": "2020-12-31T14:35:31.737Z",
          "content": "<p>I personnaly use one dictionnary per user_id. <br>\nI load them one by one during inference depending on ids seen in the test set. Overwise, my cache dictionnary would not fit.</p>",
          "rawMarkdown": "I personnaly use one dictionnary per user_id. \nI load them one by one during inference depending on ids seen in the test set. Overwise, my cache dictionnary would not fit."
        },
        {
          "id": 1133805,
          "postDate": "2020-12-31T14:43:03.637Z",
          "content": "<p>My code has a lot to be optimized, but I'm not good enough.I hope I can get 0.788 after 4hours inference</p>",
          "rawMarkdown": "My code has a lot to be optimized, but I'm not good enough.I hope I can get 0.788 after 4hours inference"
        },
        {
          "id": 1133871,
          "postDate": "2020-12-31T15:38:54.750Z",
          "content": "<p>SQLite is easy to learn(sql is pandas-like). Moreover, you need only insert(update on conflict in key, for example, on conflict in user_content_id, insert into yourtable values Your_values on conflict(some_id) do update set yourfeature=your_vslue) and select</p>",
          "rawMarkdown": "SQLite is easy to learn(sql is pandas-like). Moreover, you need only insert(update on conflict in key, for example, on conflict in user_content_id, insert into yourtable values Your_values on conflict(some_id) do update set yourfeature=your_vslue) and select\n"
        },
        {
          "id": 1134026,
          "postDate": "2020-12-31T18:51:23.617Z",
          "content": "<p>thank you , I will to use , now I need 11G when inference, too large to accept</p>",
          "rawMarkdown": "thank you , I will to use , now I need 11G when inference, too large to accept"
        }
      ]
    },
    {
      "id": 1133667,
      "postDate": "2020-12-31T12:18:01.263Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1133692,
          "postDate": "2020-12-31T12:45:48.347Z",
          "content": "<p>dict(dict)</p>",
          "rawMarkdown": "dict(dict)"
        },
        {
          "id": 1133773,
          "postDate": "2020-12-31T14:15:55.613Z",
          "content": "<p>I use joblib.dump(dict(),filename)</p>",
          "rawMarkdown": "I use joblib.dump(dict(),filename)",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1133513,
      "author_name": "Abdessalem Boukil",
      "author_url": "",
      "post_date": "2020-12-31T09:24:47.087000",
      "content": "<p>You can use pickle to dump a dictionary, like this: </p>\n<pre><code>with open('dict', 'wb') as handle:\n    pickle.dump(dict, handle, protocol=pickle.HIGHEST_PROTOCOL)\n</code></pre>\n<p>If memory overflow still persists, you can use <a href=\"https://pypi.org/project/hickle/\" target=\"_blank\">hickle</a> (HDF5 alternative to pickle). It worked for me when the dict is too large, but it takes an awfully a long time to dump the dictionary.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1133529,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-31T09:36:47.627000",
          "content": "<p>thank you,attempt dict is too large..</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1133693,
      "author_name": "Alyona Pasevieva",
      "author_url": "",
      "post_date": "2020-12-31T12:46:23.937000",
      "content": "<p>I use SQLite for nested dictionaries and pickle.dump /pickle.load for others. Pretty fast and it take only 1gb ram, when I load prepared features to kernel</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1133697,
          "author_name": "Alyona Pasevieva",
          "author_url": "",
          "post_date": "2020-12-31T12:49:31.607000",
          "content": "<p>P.S. Nested dictionaries take up most of the memory</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133713,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-31T13:10:02.903000",
          "content": "<p>yes,nested dict use too much memory,Is sqlite fast to learn?I have never used this</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133714,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-31T13:11:36.003000",
          "content": "<p>Double nested dictionaries are just too big, normal nested dictionaries don't consume too much memory</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133776,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-31T14:18:08.050000",
          "content": "<p>It looks like sql is the optimal solution for me, and I'm already consuming 11 G of memory in my inference notebook.. I can't emsemble saint model…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133798,
          "author_name": "Jacky",
          "author_url": "",
          "post_date": "2020-12-31T14:35:31.737000",
          "content": "<p>I personnaly use one dictionnary per user_id. <br>\nI load them one by one during inference depending on ids seen in the test set. Overwise, my cache dictionnary would not fit.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133805,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-31T14:43:03.637000",
          "content": "<p>My code has a lot to be optimized, but I'm not good enough.I hope I can get 0.788 after 4hours inference</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133871,
          "author_name": "Alyona Pasevieva",
          "author_url": "",
          "post_date": "2020-12-31T15:38:54.750000",
          "content": "<p>SQLite is easy to learn(sql is pandas-like). Moreover, you need only insert(update on conflict in key, for example, on conflict in user_content_id, insert into yourtable values Your_values on conflict(some_id) do update set yourfeature=your_vslue) and select</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1134026,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-31T18:51:23.617000",
          "content": "<p>thank you , I will to use , now I need 11G when inference, too large to accept</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1133667,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-31T12:18:01.263000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1133692,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-31T12:45:48.347000",
          "content": "<p>dict(dict)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133773,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-31T14:15:55.613000",
          "content": "<p>I use joblib.dump(dict(),filename)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1133488": "What format is normally used when using dump to compress dictionaries?\nI converted the 6G dictionary to zip format memory overflow...\nanother fomat can I choose?",
    "1133513": "You can use pickle to dump a dictionary, like this: \n\n    with open('dict', 'wb') as handle:\n        pickle.dump(dict, handle, protocol=pickle.HIGHEST_PROTOCOL)\n\nIf memory overflow still persists, you can use [hickle](https://pypi.org/project/hickle/) (HDF5 alternative to pickle). It worked for me when the dict is too large, but it takes an awfully a long time to dump the dictionary.",
    "1133693": "I use SQLite for nested dictionaries and pickle.dump /pickle.load for others. Pretty fast and it take only 1gb ram, when I load prepared features to kernel",
    "1133667": ""
  }
}