{
  "id": 202965,
  "title": "How to handle user-content-wise data efficiently?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/202965",
  "author_name": "",
  "post_date": "2020-12-12T23:42:18.764019600Z",
  "votes": 10,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Many people would think that user-content-wise data (i.e., mean/sum/count of accuracy/time/explanation/lecture for each user_id and for each content_id) are very effective for prediction.<br>\nI currently handle these data with nested dict (user-wise dict of content-wise dict) like following.</p>\n<pre><code>for row in df:\n   aaa_uc_dict[user_id][content_id] = aaa\n   bbb_uc_dict[user_id][content_id] = bbb\n   ccc_uc_dict[user_id][content_id] = ccc\n</code></pre>\n<p>I have created 15 dict files of user-content-wise data with my local PC. The size of these dict files become more than 2GB/file, when I saved them in pickle format. I want to use these files in kaggle notebook for inference in prediction loop, but it would easily cause memory error. </p>\n<p>Does anyone have ideas for handling these data more efficiently, or I must select two or three parameters?</p>",
  "messages": [
    {
      "id": "1110656",
      "postDate": "12/12/2020 23:42:18",
      "content": "<p>Many people would think that user-content-wise data (i.e., mean/sum/count of accuracy/time/explanation/lecture for each user_id and for each content_id) are very effective for prediction.<br>\nI currently handle these data with nested dict (user-wise dict of content-wise dict) like following.</p>\n<pre><code>for row in df:\n   aaa_uc_dict[user_id][content_id] = aaa\n   bbb_uc_dict[user_id][content_id] = bbb\n   ccc_uc_dict[user_id][content_id] = ccc\n</code></pre>\n<p>I have created 15 dict files of user-content-wise data with my local PC. The size of these dict files become more than 2GB/file, when I saved them in pickle format. I want to use these files in kaggle notebook for inference in prediction loop, but it would easily cause memory error. </p>\n<p>Does anyone have ideas for handling these data more efficiently, or I must select two or three parameters?</p>",
      "rawMarkdown": "Many people would think that user-content-wise data (i.e., mean/sum/count of accuracy/time/explanation/lecture for each user_id and for each content_id) are very effective for prediction.\nI currently handle these data with nested dict (user-wise dict of content-wise dict) like following.\n\n```\nfor row in df:\n   aaa_uc_dict[user_id][content_id] = aaa\n   bbb_uc_dict[user_id][content_id] = bbb\n   ccc_uc_dict[user_id][content_id] = ccc\n```\n\nI have created 15 dict files of user-content-wise data with my local PC. The size of these dict files become more than 2GB/file, when I saved them in pickle format. I want to use these files in kaggle notebook for inference in prediction loop, but it would easily cause memory error. \n\nDoes anyone have ideas for handling these data more efficiently, or I must select two or three parameters?",
      "votes": null
    },
    {
      "id": "1111686",
      "postDate": "12/13/2020 23:13:55",
      "content": "<p>user_id is int16 and if you use a copy of it in each dictionary it will consume the memory, try to use it once(make dictionary of dictionaries). for example you have a dictionary of users where each user have a dictionary of questions. <br>\nI used that in my kaggle notebook and it saved me huge amount of memory.</p>",
      "rawMarkdown": "user_id is int16 and if you use a copy of it in each dictionary it will consume the memory, try to use it once(make dictionary of dictionaries). for example you have a dictionary of users where each user have a dictionary of questions. \nI used that in my kaggle notebook and it saved me huge amount of memory.",
      "votes": null
    },
    {
      "id": "1111723",
      "postDate": "12/14/2020 00:50:51",
      "content": "<p>Thank you for your comment.<br>\nI have created double nested dictionaries (e.g., aaa_uc_dict[user_id][content_id]: user-wise dictionary of content-wise dictionaries for parameter aaa), but it cost me 2GB/file. <br>\nTriple nested dictionary (e.g., uc_dict[parameter_type][user_id][content_id]: parameter-wise dictionary of user-wise dictionaries of content-wise dictionaries) did not work well, as a size of single file became too huge to save and load.</p>",
      "rawMarkdown": "Thank you for your comment.\nI have created double nested dictionaries (e.g., aaa_uc_dict[user_id][content_id]: user-wise dictionary of content-wise dictionaries for parameter aaa), but it cost me 2GB/file. \nTriple nested dictionary (e.g., uc_dict[parameter_type][user_id][content_id]: parameter-wise dictionary of user-wise dictionaries of content-wise dictionaries) did not work well, as a size of single file became too huge to save and load.",
      "votes": null
    },
    {
      "id": "1112877",
      "postDate": "12/15/2020 00:55:42",
      "content": "<p>I . <br>\nIndices of dict were assigned as np.int in my script. Transforming values into int reduced the size of dict into 1GB for float parameters and 400MB for int parameters.</p>\n<pre><code>for cnt row in df[\"user_id\", \"content_id\",\"aaa\"].values:\n   user_id=int(row[0])\n   content_id=int(row[1])\n   aaa=int(row[2])\n   aaa_uc_dict[user_id][content_id] = aaa\n</code></pre>\n<p>Any other ideas are welcome.</p>",
      "rawMarkdown": "I ~~self-solved this problem partially~~. \nIndices of dict were assigned as np.int in my script. Transforming values into int reduced the size of dict into 1GB for float parameters and 400MB for int parameters.\n\n```\nfor cnt row in df[\"user_id\", \"content_id\",\"aaa\"].values:\n   user_id=int(row[0])\n   content_id=int(row[1])\n   aaa=int(row[2])\n   aaa_uc_dict[user_id][content_id] = aaa\n```\n\nAny other ideas are welcome.",
      "votes": null
    },
    {
      "id": "1119876",
      "postDate": "12/20/2020 12:24:10",
      "content": "<p>I noticed that my 428 MB dict file (total count of interactions for each user_id and each content_id) consumes 5.9 GB memory (measured with <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203020\" target=\"_blank\">this great script</a> by <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a>).<br>\nI am still struggling with this issue.<br>\nAny solutions or comments are highly welcome.</p>",
      "rawMarkdown": "I noticed that my 428 MB dict file (total count of interactions for each user_id and each content_id) consumes 5.9 GB memory (measured with [this great script](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203020) by @higepon).\nI am still struggling with this issue.\nAny solutions or comments are highly welcome.",
      "votes": null
    },
    {
      "id": "1120202",
      "postDate": "12/20/2020 16:57:56",
      "content": "<p>Yes not easy.. I think it should be something where not all the data has to be loaded into memory..</p>\n<p>I tried:</p>\n<ul>\n<li>sqlite: too slow to query the data</li>\n<li>h5py with users as groups and contents as subgroups: too large and too slow</li>\n<li>h5py with just users as groups and the content-data as numpy array: seems to work quite well in terms of size and speed but indexing and averaging seems cumbersome..</li>\n</ul>",
      "rawMarkdown": "Yes not easy.. I think it should be something where not all the data has to be loaded into memory..\n\nI tried:\n- sqlite: too slow to query the data\n- h5py with users as groups and contents as subgroups: too large and too slow\n- h5py with just users as groups and the content-data as numpy array: seems to work quite well in terms of size and speed but indexing and averaging seems cumbersome..",
      "votes": null
    },
    {
      "id": "1120643",
      "postDate": "12/21/2020 02:25:00",
      "content": "<p>Thank you for your comment.<br>\nWould you explain more about your h5py data structure?<br>\nFor example, do you create an array of np.zeros(len(questions)) for each user at first?</p>",
      "rawMarkdown": "Thank you for your comment.\nWould you explain more about your h5py data structure?\nFor example, do you create an array of np.zeros(len(questions)) for each user at first?",
      "votes": null
    },
    {
      "id": "1120918",
      "postDate": "12/21/2020 07:47:23",
      "content": "<p>The issue is that numbers inside dicts are store as int or float <strong>objects</strong> which consumes aboute 20+ bytes for each number (while in numpy it is maximum 4 bytes!). So if you have dict of dicts, then the number of bytes will be enormous. <br>\nActually I didn't find any solution for this problem, so I left the whole idea of creating dict of dicts !<br>\nIf you find any idea please let me know</p>",
      "rawMarkdown": "The issue is that numbers inside dicts are store as int or float **objects** which consumes aboute 20+ bytes for each number (while in numpy it is maximum 4 bytes!). So if you have dict of dicts, then the number of bytes will be enormous. \nActually I didn't find any solution for this problem, so I left the whole idea of creating dict of dicts !\nIf you find any idea please let me know",
      "votes": null
    },
    {
      "id": "1121311",
      "postDate": "12/21/2020 14:39:00",
      "content": "<p>Still in the process of figuring out the best approach.. but so far: first defined each user as a group member/\"key\" and then created an array (as dataset) for each user with content_id as the first column and the other variables in the other columns. Not sure yet how well it works..</p>",
      "rawMarkdown": "Still in the process of figuring out the best approach.. but so far: first defined each user as a group member/\"key\" and then created an array (as dataset) for each user with content_id as the first column and the other variables in the other columns. Not sure yet how well it works..",
      "votes": null
    },
    {
      "id": "1121825",
      "postDate": "12/21/2020 23:59:41",
      "content": "<p>I had a similar issue and ended up using sqlite3.<br>\nHere are what I'm doing :) Hope it helps.</p>\n<h3>For training</h3>\n<p>I have a nested defaultdict such as foo[user_id][content_id] for feature engineering.<br>\nThis consumes a lot of memory but which is fine because I'm training on GCP. I just increased VM memory.<br>\nOnce training is done. I save the nested defaultdict into sqlite3 table.<br>\nThe table looks like below</p>\n<pre><code>CREATE TABLE foot(user_id INTEGER, content_id INTEGER,  foo INTEGER, PRIMARY KEY (user_id, content_id))\n</code></pre>\n<p>You'll have .db file for the database.</p>\n<h3>For submission</h3>\n<p>Because Kaggle kernel doesn't have enough memory to have the nested dic in memory I \"SELECT/INSERT/UPDATE\" the foo table whenever it's necessary for submission. The db is on disk so it doesn't consume much memory.</p>\n<h3>Speed?</h3>\n<p>Someone mentioned sqlite3 is slow which is true. But it's fast enough for <em>submission</em>.<br>\nIf you think sqlite3 is not fast enough you can take a look at the following great article.<br>\n<a href=\"https://stackoverflow.com/questions/1711631/improve-insert-per-second-performance-of-sqlite\" target=\"_blank\">https://stackoverflow.com/questions/1711631/improve-insert-per-second-performance-of-sqlite</a></p>",
      "rawMarkdown": "I had a similar issue and ended up using sqlite3.\nHere are what I'm doing :) Hope it helps.\n\n### For training\nI have a nested defaultdict such as foo[user_id][content_id] for feature engineering.\nThis consumes a lot of memory but which is fine because I'm training on GCP. I just increased VM memory.\nOnce training is done. I save the nested defaultdict into sqlite3 table.\nThe table looks like below\n```\nCREATE TABLE foot(user_id INTEGER, content_id INTEGER,  foo INTEGER, PRIMARY KEY (user_id, content_id))\n```\nYou'll have .db file for the database.\n\n### For submission\nBecause Kaggle kernel doesn't have enough memory to have the nested dic in memory I \"SELECT/INSERT/UPDATE\" the foo table whenever it's necessary for submission. The db is on disk so it doesn't consume much memory.\n\n### Speed?\nSomeone mentioned sqlite3 is slow which is true. But it's fast enough for *submission*.\nIf you think sqlite3 is not fast enough you can take a look at the following great article.\nhttps://stackoverflow.com/questions/1711631/improve-insert-per-second-performance-of-sqlite",
      "votes": null
    },
    {
      "id": "1121845",
      "postDate": "12/22/2020 00:33:05",
      "content": "<p>Thank you for providing valuable information and code!<br>\nI will try sqlite3.</p>",
      "rawMarkdown": "Thank you for providing valuable information and code!\nI will try sqlite3.",
      "votes": null
    },
    {
      "id": "1124163",
      "postDate": "12/23/2020 17:52:17",
      "content": "<p>I was finding sqlite working quite quickly. <a href=\"https://www.kaggle.com/calebeverett/riiid-submit\" target=\"_blank\">This book</a> submits in under three hours with lookup and update of user content. I found that the insert wasn't the problem, it was the merging the selected data with the batch from the test iterator. The bit below, I think is what makes it fast, where you essentially bring the values from the batch into the select query and let sqlite do the merge and then deliver a complete batch ready for prediction, avoiding any pandas merge/join altogether.</p>\n<pre><code>records = df_batch[batch_cols].fillna(0).to_records(index=False)\n\ndef select_state(batch_cols, records):\n    return f\"\"\"\n        WITH b ({(', ').join(batch_cols)}) AS (\n        VALUES {(', ').join(list(map(str, records)))}\n        )\n        SELECT\n            {(', ').join([f'b.{col}' for col in batch_cols])},\n            IFNULL(answered_correctly_cumsum, 0) answered_correctly_cumsum, \n...\n        FROM b\n        LEFT JOIN (\n            SELECT user_id, answered_correctly answered_correctly_cumsum,\n                answered_incorrectly answered_incorrectly_cumsum\n            FROM users\n            WHERE {(' OR ').join([f'user_id = {r[0]}' for r in records])}\n        ) u ON (u.user_id = b.user_id)\n...\n</code></pre>\n<p>The other trick for the update was to use the on conflict clause to update existing values or insert new ones.</p>",
      "rawMarkdown": "I was finding sqlite working quite quickly. [This book](https://www.kaggle.com/calebeverett/riiid-submit) submits in under three hours with lookup and update of user content. I found that the insert wasn't the problem, it was the merging the selected data with the batch from the test iterator. The bit below, I think is what makes it fast, where you essentially bring the values from the batch into the select query and let sqlite do the merge and then deliver a complete batch ready for prediction, avoiding any pandas merge/join altogether.\n\n```\nrecords = df_batch[batch_cols].fillna(0).to_records(index=False)\n\ndef select_state(batch_cols, records):\n    return f\"\"\"\n        WITH b ({(', ').join(batch_cols)}) AS (\n        VALUES {(', ').join(list(map(str, records)))}\n        )\n        SELECT\n            {(', ').join([f'b.{col}' for col in batch_cols])},\n            IFNULL(answered_correctly_cumsum, 0) answered_correctly_cumsum, \n...\n        FROM b\n        LEFT JOIN (\n            SELECT user_id, answered_correctly answered_correctly_cumsum,\n                answered_incorrectly answered_incorrectly_cumsum\n            FROM users\n            WHERE {(' OR ').join([f'user_id = {r[0]}' for r in records])}\n        ) u ON (u.user_id = b.user_id)\n...\n```\n\nThe other trick for the update was to use the on conflict clause to update existing values or insert new ones.",
      "votes": null
    },
    {
      "id": "1124519",
      "postDate": "12/24/2020 02:40:00",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a>. <br>\nI have read your notebooks already, but did not notice its potential (because I could not understand bigquery or sqlite).  Now I realize how helpful it is.  I will learn many techniques in your notebook.</p>",
      "rawMarkdown": "Thank you @calebeverett. \nI have read your notebooks already, but did not notice its potential (because I could not understand bigquery or sqlite).  Now I realize how helpful it is.  I will learn many techniques in your notebook.",
      "votes": null
    },
    {
      "id": "1125288",
      "postDate": "12/24/2020 15:37:34",
      "content": "<p>That would be great if you found something useful there. I think you could load your dataframes into sqlite for the predictions from your current feature engineering methodology without having to spend time understanding the other big query book.</p>",
      "rawMarkdown": "That would be great if you found something useful there. I think you could load your dataframes into sqlite for the predictions from your current feature engineering methodology without having to spend time understanding the other big query book.",
      "votes": null
    },
    {
      "id": "1125493",
      "postDate": "12/24/2020 18:53:40",
      "content": "<p>If you're not afraid of having a bit of noise, there's a way to store data in a restricted amount of memory.</p>\n<p>Choose a big Dimension value:<br>\nFor instance D=20_000_000<br>\nThen create a list or a np.array:<br>\nMyList=[0]*D<br>\nMyArray=np.zeros(D)<br>\nThat's all the amount of memory you'll need.</p>\n<p>Then you can store anything inside with a chance of collision accessing items like this : <br>\nMyArray[abs(hash('any generated key')) % D]</p>\n<p>This way you don't have to store index data, and you'll never use more memory than you have.<br>\nThe problems are : </p>\n<ul>\n<li>The more data you store, the more the risk of collision. (and we can't estimate the impact of new data)</li>\n<li>The computation of abs(hash('any generated key')) % D takes time.</li>\n</ul>",
      "rawMarkdown": "If you're not afraid of having a bit of noise, there's a way to store data in a restricted amount of memory.\n\nChoose a big Dimension value:\nFor instance D=20_000_000\nThen create a list or a np.array:\nMyList=[0]*D\nMyArray=np.zeros(D)\nThat's all the amount of memory you'll need.\n\nThen you can store anything inside with a chance of collision accessing items like this : \nMyArray[abs(hash('any generated key')) % D]\n\nThis way you don't have to store index data, and you'll never use more memory than you have.\nThe problems are : \n- The more data you store, the more the risk of collision. (and we can't estimate the impact of new data)\n- The computation of abs(hash('any generated key')) % D takes time.",
      "votes": null
    },
    {
      "id": "1128936",
      "postDate": "12/27/2020 22:42:10",
      "content": "<p><a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> Hi! Can you tell me please, if you load .db file to Kaggle dataset? I have an error, when try to do in this way</p>",
      "rawMarkdown": "higepon @calebeverett Hi! Can you tell me please, if you load .db file to Kaggle dataset? I have an error, when try to do in this way",
      "votes": null
    },
    {
      "id": "1128953",
      "postDate": "12/27/2020 23:22:38",
      "content": "<p>The following is working for me. Note that I'm copying the db file to current directory which is required when you write.</p>\n<pre><code>data_dir = Path('/kaggle/input/higeponriiid2')\n!cp /kaggle/input/higeponriiid2/riiid.db riiid.db\nconn = sqlite3.connect('riiid.db')\n</code></pre>\n<p>Hope it helps.</p>",
      "rawMarkdown": "The following is working for me. Note that I'm copying the db file to current directory which is required when you write.\n\n```\ndata_dir = Path('/kaggle/input/higeponriiid2')\n!cp /kaggle/input/higeponriiid2/riiid.db riiid.db\nconn = sqlite3.connect('riiid.db')\n```\n\nHope it helps.",
      "votes": null
    },
    {
      "id": "1128954",
      "postDate": "12/27/2020 23:24:42",
      "content": "<p>Forgot to mention but I upload the db to kaggle dataset which is higeponriid2 above.</p>",
      "rawMarkdown": "Forgot to mention but I upload the db to kaggle dataset which is higeponriid2 above.",
      "votes": null
    },
    {
      "id": "1128955",
      "postDate": "12/27/2020 23:27:30",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> <br>\nI’ll try it!</p>",
      "rawMarkdown": "Thank you @higepon \nI’ll try it!",
      "votes": null
    },
    {
      "id": "1129038",
      "postDate": "12/28/2020 02:57:27",
      "content": "<p>I was using data frames and loading to db in submit book. They are in the dataset attached to the submit book. The database should be available in the book after you run.</p>",
      "rawMarkdown": "I was using data frames and loading to db in submit book. They are in the dataset attached to the submit book. The database should be available in the book after you run.",
      "votes": null
    },
    {
      "id": "1129127",
      "postDate": "12/28/2020 05:23:00",
      "content": "<p>Try using a dataframe that shares an index. Both speed and memory are fine.<br>\nIf there is one key, a dict is used, and if there are two keys, a dataframe is used.</p>\n<p>Full dataset<br>\nmemory: ~7gb<br>\ntime: ~4hour</p>",
      "rawMarkdown": "Try using a dataframe that shares an index. Both speed and memory are fine.\nIf there is one key, a dict is used, and if there are two keys, a dataframe is used.\n\nFull dataset\nmemory: ~7gb\ntime: ~4hour",
      "votes": null
    },
    {
      "id": "1129629",
      "postDate": "12/28/2020 13:26:31",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> </p>",
      "rawMarkdown": "Thank you @calebeverett",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1111686,
      "author_name": "mohamadnawfal",
      "author_url": "",
      "post_date": "12/13/2020 23:13:55",
      "content": "<p>user_id is int16 and if you use a copy of it in each dictionary it will consume the memory, try to use it once(make dictionary of dictionaries). for example you have a dictionary of users where each user have a dictionary of questions. <br>\nI used that in my kaggle notebook and it saved me huge amount of memory.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1111723,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "12/14/2020 00:50:51",
          "content": "<p>Thank you for your comment.<br>\nI have created double nested dictionaries (e.g., aaa_uc_dict[user_id][content_id]: user-wise dictionary of content-wise dictionaries for parameter aaa), but it cost me 2GB/file. <br>\nTriple nested dictionary (e.g., uc_dict[parameter_type][user_id][content_id]: parameter-wise dictionary of user-wise dictionaries of content-wise dictionaries) did not work well, as a size of single file became too huge to save and load.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1112877,
      "author_name": "tomooinubushi",
      "author_url": "",
      "post_date": "12/15/2020 00:55:42",
      "content": "<p>I . <br>\nIndices of dict were assigned as np.int in my script. Transforming values into int reduced the size of dict into 1GB for float parameters and 400MB for int parameters.</p>\n<pre><code>for cnt row in df[\"user_id\", \"content_id\",\"aaa\"].values:\n   user_id=int(row[0])\n   content_id=int(row[1])\n   aaa=int(row[2])\n   aaa_uc_dict[user_id][content_id] = aaa\n</code></pre>\n<p>Any other ideas are welcome.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1119876,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "12/20/2020 12:24:10",
          "content": "<p>I noticed that my 428 MB dict file (total count of interactions for each user_id and each content_id) consumes 5.9 GB memory (measured with <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203020\" target=\"_blank\">this great script</a> by <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a>).<br>\nI am still struggling with this issue.<br>\nAny solutions or comments are highly welcome.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1120918,
          "author_name": "mohamadnawfal",
          "author_url": "",
          "post_date": "12/21/2020 07:47:23",
          "content": "<p>The issue is that numbers inside dicts are store as int or float <strong>objects</strong> which consumes aboute 20+ bytes for each number (while in numpy it is maximum 4 bytes!). So if you have dict of dicts, then the number of bytes will be enormous. <br>\nActually I didn't find any solution for this problem, so I left the whole idea of creating dict of dicts !<br>\nIf you find any idea please let me know</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121825,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/21/2020 23:59:41",
          "content": "<p>I had a similar issue and ended up using sqlite3.<br>\nHere are what I'm doing :) Hope it helps.</p>\n<h3>For training</h3>\n<p>I have a nested defaultdict such as foo[user_id][content_id] for feature engineering.<br>\nThis consumes a lot of memory but which is fine because I'm training on GCP. I just increased VM memory.<br>\nOnce training is done. I save the nested defaultdict into sqlite3 table.<br>\nThe table looks like below</p>\n<pre><code>CREATE TABLE foot(user_id INTEGER, content_id INTEGER,  foo INTEGER, PRIMARY KEY (user_id, content_id))\n</code></pre>\n<p>You'll have .db file for the database.</p>\n<h3>For submission</h3>\n<p>Because Kaggle kernel doesn't have enough memory to have the nested dic in memory I \"SELECT/INSERT/UPDATE\" the foo table whenever it's necessary for submission. The db is on disk so it doesn't consume much memory.</p>\n<h3>Speed?</h3>\n<p>Someone mentioned sqlite3 is slow which is true. But it's fast enough for <em>submission</em>.<br>\nIf you think sqlite3 is not fast enough you can take a look at the following great article.<br>\n<a href=\"https://stackoverflow.com/questions/1711631/improve-insert-per-second-performance-of-sqlite\" target=\"_blank\">https://stackoverflow.com/questions/1711631/improve-insert-per-second-performance-of-sqlite</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121845,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "12/22/2020 00:33:05",
          "content": "<p>Thank you for providing valuable information and code!<br>\nI will try sqlite3.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124163,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "12/23/2020 17:52:17",
          "content": "<p>I was finding sqlite working quite quickly. <a href=\"https://www.kaggle.com/calebeverett/riiid-submit\" target=\"_blank\">This book</a> submits in under three hours with lookup and update of user content. I found that the insert wasn't the problem, it was the merging the selected data with the batch from the test iterator. The bit below, I think is what makes it fast, where you essentially bring the values from the batch into the select query and let sqlite do the merge and then deliver a complete batch ready for prediction, avoiding any pandas merge/join altogether.</p>\n<pre><code>records = df_batch[batch_cols].fillna(0).to_records(index=False)\n\ndef select_state(batch_cols, records):\n    return f\"\"\"\n        WITH b ({(', ').join(batch_cols)}) AS (\n        VALUES {(', ').join(list(map(str, records)))}\n        )\n        SELECT\n            {(', ').join([f'b.{col}' for col in batch_cols])},\n            IFNULL(answered_correctly_cumsum, 0) answered_correctly_cumsum, \n...\n        FROM b\n        LEFT JOIN (\n            SELECT user_id, answered_correctly answered_correctly_cumsum,\n                answered_incorrectly answered_incorrectly_cumsum\n            FROM users\n            WHERE {(' OR ').join([f'user_id = {r[0]}' for r in records])}\n        ) u ON (u.user_id = b.user_id)\n...\n</code></pre>\n<p>The other trick for the update was to use the on conflict clause to update existing values or insert new ones.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124519,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "12/24/2020 02:40:00",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a>. <br>\nI have read your notebooks already, but did not notice its potential (because I could not understand bigquery or sqlite).  Now I realize how helpful it is.  I will learn many techniques in your notebook.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125288,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "12/24/2020 15:37:34",
          "content": "<p>That would be great if you found something useful there. I think you could load your dataframes into sqlite for the predictions from your current feature engineering methodology without having to spend time understanding the other big query book.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1128936,
          "author_name": "fredegrec",
          "author_url": "",
          "post_date": "12/27/2020 22:42:10",
          "content": "<p><a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> Hi! Can you tell me please, if you load .db file to Kaggle dataset? I have an error, when try to do in this way</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1128953,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/27/2020 23:22:38",
          "content": "<p>The following is working for me. Note that I'm copying the db file to current directory which is required when you write.</p>\n<pre><code>data_dir = Path('/kaggle/input/higeponriiid2')\n!cp /kaggle/input/higeponriiid2/riiid.db riiid.db\nconn = sqlite3.connect('riiid.db')\n</code></pre>\n<p>Hope it helps.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1128954,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/27/2020 23:24:42",
          "content": "<p>Forgot to mention but I upload the db to kaggle dataset which is higeponriid2 above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1128955,
          "author_name": "fredegrec",
          "author_url": "",
          "post_date": "12/27/2020 23:27:30",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> <br>\nI’ll try it!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1129038,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "12/28/2020 02:57:27",
          "content": "<p>I was using data frames and loading to db in submit book. They are in the dataset attached to the submit book. The database should be available in the book after you run.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1129629,
          "author_name": "fredegrec",
          "author_url": "",
          "post_date": "12/28/2020 13:26:31",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1120202,
      "author_name": "hannes82",
      "author_url": "",
      "post_date": "12/20/2020 16:57:56",
      "content": "<p>Yes not easy.. I think it should be something where not all the data has to be loaded into memory..</p>\n<p>I tried:</p>\n<ul>\n<li>sqlite: too slow to query the data</li>\n<li>h5py with users as groups and contents as subgroups: too large and too slow</li>\n<li>h5py with just users as groups and the content-data as numpy array: seems to work quite well in terms of size and speed but indexing and averaging seems cumbersome..</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1120643,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "12/21/2020 02:25:00",
          "content": "<p>Thank you for your comment.<br>\nWould you explain more about your h5py data structure?<br>\nFor example, do you create an array of np.zeros(len(questions)) for each user at first?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121311,
          "author_name": "hannes82",
          "author_url": "",
          "post_date": "12/21/2020 14:39:00",
          "content": "<p>Still in the process of figuring out the best approach.. but so far: first defined each user as a group member/\"key\" and then created an array (as dataset) for each user with content_id as the first column and the other variables in the other columns. Not sure yet how well it works..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1125493,
      "author_name": "trottefox",
      "author_url": "",
      "post_date": "12/24/2020 18:53:40",
      "content": "<p>If you're not afraid of having a bit of noise, there's a way to store data in a restricted amount of memory.</p>\n<p>Choose a big Dimension value:<br>\nFor instance D=20_000_000<br>\nThen create a list or a np.array:<br>\nMyList=[0]*D<br>\nMyArray=np.zeros(D)<br>\nThat's all the amount of memory you'll need.</p>\n<p>Then you can store anything inside with a chance of collision accessing items like this : <br>\nMyArray[abs(hash('any generated key')) % D]</p>\n<p>This way you don't have to store index data, and you'll never use more memory than you have.<br>\nThe problems are : </p>\n<ul>\n<li>The more data you store, the more the risk of collision. (and we can't estimate the impact of new data)</li>\n<li>The computation of abs(hash('any generated key')) % D takes time.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1129127,
      "author_name": "yeonmin",
      "author_url": "",
      "post_date": "12/28/2020 05:23:00",
      "content": "<p>Try using a dataframe that shares an index. Both speed and memory are fine.<br>\nIf there is one key, a dict is used, and if there are two keys, a dataframe is used.</p>\n<p>Full dataset<br>\nmemory: ~7gb<br>\ntime: ~4hour</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1110656": "Many people would think that user-content-wise data (i.e., mean/sum/count of accuracy/time/explanation/lecture for each user_id and for each content_id) are very effective for prediction.\nI currently handle these data with nested dict (user-wise dict of content-wise dict) like following.\n\n```\nfor row in df:\n   aaa_uc_dict[user_id][content_id] = aaa\n   bbb_uc_dict[user_id][content_id] = bbb\n   ccc_uc_dict[user_id][content_id] = ccc\n```\n\nI have created 15 dict files of user-content-wise data with my local PC. The size of these dict files become more than 2GB/file, when I saved them in pickle format. I want to use these files in kaggle notebook for inference in prediction loop, but it would easily cause memory error. \n\nDoes anyone have ideas for handling these data more efficiently, or I must select two or three parameters?",
    "1111686": "user_id is int16 and if you use a copy of it in each dictionary it will consume the memory, try to use it once(make dictionary of dictionaries). for example you have a dictionary of users where each user have a dictionary of questions. \nI used that in my kaggle notebook and it saved me huge amount of memory.",
    "1111723": "Thank you for your comment.\nI have created double nested dictionaries (e.g., aaa_uc_dict[user_id][content_id]: user-wise dictionary of content-wise dictionaries for parameter aaa), but it cost me 2GB/file. \nTriple nested dictionary (e.g., uc_dict[parameter_type][user_id][content_id]: parameter-wise dictionary of user-wise dictionaries of content-wise dictionaries) did not work well, as a size of single file became too huge to save and load.",
    "1112877": "I ~~self-solved this problem partially~~. \nIndices of dict were assigned as np.int in my script. Transforming values into int reduced the size of dict into 1GB for float parameters and 400MB for int parameters.\n\n```\nfor cnt row in df[\"user_id\", \"content_id\",\"aaa\"].values:\n   user_id=int(row[0])\n   content_id=int(row[1])\n   aaa=int(row[2])\n   aaa_uc_dict[user_id][content_id] = aaa\n```\n\nAny other ideas are welcome.",
    "1119876": "I noticed that my 428 MB dict file (total count of interactions for each user_id and each content_id) consumes 5.9 GB memory (measured with [this great script](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203020) by @higepon).\nI am still struggling with this issue.\nAny solutions or comments are highly welcome.",
    "1120202": "Yes not easy.. I think it should be something where not all the data has to be loaded into memory..\n\nI tried:\n- sqlite: too slow to query the data\n- h5py with users as groups and contents as subgroups: too large and too slow\n- h5py with just users as groups and the content-data as numpy array: seems to work quite well in terms of size and speed but indexing and averaging seems cumbersome..",
    "1120643": "Thank you for your comment.\nWould you explain more about your h5py data structure?\nFor example, do you create an array of np.zeros(len(questions)) for each user at first?",
    "1120918": "The issue is that numbers inside dicts are store as int or float **objects** which consumes aboute 20+ bytes for each number (while in numpy it is maximum 4 bytes!). So if you have dict of dicts, then the number of bytes will be enormous. \nActually I didn't find any solution for this problem, so I left the whole idea of creating dict of dicts !\nIf you find any idea please let me know",
    "1121311": "Still in the process of figuring out the best approach.. but so far: first defined each user as a group member/\"key\" and then created an array (as dataset) for each user with content_id as the first column and the other variables in the other columns. Not sure yet how well it works..",
    "1121825": "I had a similar issue and ended up using sqlite3.\nHere are what I'm doing :) Hope it helps.\n\n### For training\nI have a nested defaultdict such as foo[user_id][content_id] for feature engineering.\nThis consumes a lot of memory but which is fine because I'm training on GCP. I just increased VM memory.\nOnce training is done. I save the nested defaultdict into sqlite3 table.\nThe table looks like below\n```\nCREATE TABLE foot(user_id INTEGER, content_id INTEGER,  foo INTEGER, PRIMARY KEY (user_id, content_id))\n```\nYou'll have .db file for the database.\n\n### For submission\nBecause Kaggle kernel doesn't have enough memory to have the nested dic in memory I \"SELECT/INSERT/UPDATE\" the foo table whenever it's necessary for submission. The db is on disk so it doesn't consume much memory.\n\n### Speed?\nSomeone mentioned sqlite3 is slow which is true. But it's fast enough for *submission*.\nIf you think sqlite3 is not fast enough you can take a look at the following great article.\nhttps://stackoverflow.com/questions/1711631/improve-insert-per-second-performance-of-sqlite",
    "1121845": "Thank you for providing valuable information and code!\nI will try sqlite3.",
    "1124163": "I was finding sqlite working quite quickly. [This book](https://www.kaggle.com/calebeverett/riiid-submit) submits in under three hours with lookup and update of user content. I found that the insert wasn't the problem, it was the merging the selected data with the batch from the test iterator. The bit below, I think is what makes it fast, where you essentially bring the values from the batch into the select query and let sqlite do the merge and then deliver a complete batch ready for prediction, avoiding any pandas merge/join altogether.\n\n```\nrecords = df_batch[batch_cols].fillna(0).to_records(index=False)\n\ndef select_state(batch_cols, records):\n    return f\"\"\"\n        WITH b ({(', ').join(batch_cols)}) AS (\n        VALUES {(', ').join(list(map(str, records)))}\n        )\n        SELECT\n            {(', ').join([f'b.{col}' for col in batch_cols])},\n            IFNULL(answered_correctly_cumsum, 0) answered_correctly_cumsum, \n...\n        FROM b\n        LEFT JOIN (\n            SELECT user_id, answered_correctly answered_correctly_cumsum,\n                answered_incorrectly answered_incorrectly_cumsum\n            FROM users\n            WHERE {(' OR ').join([f'user_id = {r[0]}' for r in records])}\n        ) u ON (u.user_id = b.user_id)\n...\n```\n\nThe other trick for the update was to use the on conflict clause to update existing values or insert new ones.",
    "1124519": "Thank you @calebeverett. \nI have read your notebooks already, but did not notice its potential (because I could not understand bigquery or sqlite).  Now I realize how helpful it is.  I will learn many techniques in your notebook.",
    "1125288": "That would be great if you found something useful there. I think you could load your dataframes into sqlite for the predictions from your current feature engineering methodology without having to spend time understanding the other big query book.",
    "1125493": "If you're not afraid of having a bit of noise, there's a way to store data in a restricted amount of memory.\n\nChoose a big Dimension value:\nFor instance D=20_000_000\nThen create a list or a np.array:\nMyList=[0]*D\nMyArray=np.zeros(D)\nThat's all the amount of memory you'll need.\n\nThen you can store anything inside with a chance of collision accessing items like this : \nMyArray[abs(hash('any generated key')) % D]\n\nThis way you don't have to store index data, and you'll never use more memory than you have.\nThe problems are : \n- The more data you store, the more the risk of collision. (and we can't estimate the impact of new data)\n- The computation of abs(hash('any generated key')) % D takes time.",
    "1128936": "higepon @calebeverett Hi! Can you tell me please, if you load .db file to Kaggle dataset? I have an error, when try to do in this way",
    "1128953": "The following is working for me. Note that I'm copying the db file to current directory which is required when you write.\n\n```\ndata_dir = Path('/kaggle/input/higeponriiid2')\n!cp /kaggle/input/higeponriiid2/riiid.db riiid.db\nconn = sqlite3.connect('riiid.db')\n```\n\nHope it helps.",
    "1128954": "Forgot to mention but I upload the db to kaggle dataset which is higeponriid2 above.",
    "1128955": "Thank you @higepon \nI’ll try it!",
    "1129038": "I was using data frames and loading to db in submit book. They are in the dataset attached to the submit book. The database should be available in the book after you run.",
    "1129127": "Try using a dataframe that shares an index. Both speed and memory are fine.\nIf there is one key, a dict is used, and if there are two keys, a dataframe is used.\n\nFull dataset\nmemory: ~7gb\ntime: ~4hour",
    "1129629": "Thank you @calebeverett"
  },
  "source": "meta"
}