{
  "id": 209196,
  "title": "Dict/List Persistent Storage",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209196",
  "author_name": "",
  "post_date": "2021-01-06T16:27:11.003650300Z",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I spent quite some time trying to figure out how to decrease RAM usage. I didn't like the idea of using SQLite as the syntax is not easy and it's not easy to use dicts for that. I tried the module shelve (<a href=\"https://docs.python.org/3/library/shelve.html\" target=\"_blank\">https://docs.python.org/3/library/shelve.html</a>) which is exactly what I needed but unfortunately kaggle environnement doesn't have the backend ndbm which makes it impossible to use.<br>\nI found this repo <a href=\"https://github.com/dagnelies/pysos\" target=\"_blank\">https://github.com/dagnelies/pysos</a> which is quite awesome, it makes the job and doesn't need any backend. It's very fast and let's you use any kind of jsonable object (think dicts and lists). </p>",
  "messages": [
    {
      "id": "1141328",
      "postDate": "01/06/2021 16:27:11",
      "content": "<p>I spent quite some time trying to figure out how to decrease RAM usage. I didn't like the idea of using SQLite as the syntax is not easy and it's not easy to use dicts for that. I tried the module shelve (<a href=\"https://docs.python.org/3/library/shelve.html\" target=\"_blank\">https://docs.python.org/3/library/shelve.html</a>) which is exactly what I needed but unfortunately kaggle environnement doesn't have the backend ndbm which makes it impossible to use.<br>\nI found this repo <a href=\"https://github.com/dagnelies/pysos\" target=\"_blank\">https://github.com/dagnelies/pysos</a> which is quite awesome, it makes the job and doesn't need any backend. It's very fast and let's you use any kind of jsonable object (think dicts and lists). </p>",
      "rawMarkdown": "I spent quite some time trying to figure out how to decrease RAM usage. I didn't like the idea of using SQLite as the syntax is not easy and it's not easy to use dicts for that. I tried the module shelve (https://docs.python.org/3/library/shelve.html) which is exactly what I needed but unfortunately kaggle environnement doesn't have the backend ndbm which makes it impossible to use.\nI found this repo https://github.com/dagnelies/pysos which is quite awesome, it makes the job and doesn't need any backend. It's very fast and let's you use any kind of jsonable object (think dicts and lists).",
      "votes": null
    },
    {
      "id": "1141343",
      "postDate": "01/06/2021 16:37:36",
      "content": "<p>Be careful if you're going to use this - I tried it at the start of the competition a few months ago but found there were some bugs (unfortunately I don't remember what they were). You may still find it useful, just don't blindly trust it to solve all your problems.</p>",
      "rawMarkdown": "Be careful if you're going to use this - I tried it at the start of the competition a few months ago but found there were some bugs (unfortunately I don't remember what they were). You may still find it useful, just don't blindly trust it to solve all your problems.",
      "votes": null
    },
    {
      "id": "1141351",
      "postDate": "01/06/2021 16:43:02",
      "content": "<p>Thanks. Any alternative ?<br>\nI had some errors because I used numpy types (which are not json serializable), do you remember if it was related to that ?</p>",
      "rawMarkdown": "Thanks. Any alternative ?\nI had some errors because I used numpy types (which are not json serializable), do you remember if it was related to that ?",
      "votes": null
    },
    {
      "id": "1141391",
      "postDate": "01/06/2021 17:15:34",
      "content": "<p>I'm using SQLite. I was surprised how fast it is and it let's you maintain large statistics. With some wrapper code it almost feels like an infinite dictionary.</p>",
      "rawMarkdown": "I'm using SQLite. I was surprised how fast it is and it let's you maintain large statistics. With some wrapper code it almost feels like an infinite dictionary.",
      "votes": null
    },
    {
      "id": "1141401",
      "postDate": "01/06/2021 17:23:39",
      "content": "<p>I still tend to use sqlite, but double nested dictionaries do not know how to store with sqlite</p>",
      "rawMarkdown": "I still tend to use sqlite, but double nested dictionaries do not know how to store with sqlite",
      "votes": null
    },
    {
      "id": "1141402",
      "postDate": "01/06/2021 17:25:20",
      "content": "<p>may I know how to store double nested dictionaries with sqlite? concat two keys as one key?</p>",
      "rawMarkdown": "may I know how to store double nested dictionaries with sqlite? concat two keys as one key?",
      "votes": null
    },
    {
      "id": "1141420",
      "postDate": "01/06/2021 17:38:52",
      "content": "<p>Sure, it's even simpler: just use two columns as key.</p>\n<p>Basically like this:</p>\n<p><code>update user_question_table set key=value where user_id=... and question_id=...</code></p>\n<p>To get fast retrievals it's a good idea to also add a multi-column index on the table.</p>",
      "rawMarkdown": "Sure, it's even simpler: just use two columns as key.\n\nBasically like this:\n\n`update user_question_table set key=value where user_id=... and question_id=...`\n\nTo get fast retrievals it's a good idea to also add a multi-column index on the table.",
      "votes": null
    },
    {
      "id": "1141436",
      "postDate": "01/06/2021 17:50:22",
      "content": "<p>thank you 👍</p>",
      "rawMarkdown": "thank you 👍",
      "votes": null
    },
    {
      "id": "1141822",
      "postDate": "01/06/2021 23:19:00",
      "content": "<p>In any case someone is looking for a very easy way, my submission using <a href=\"https://github.com/dagnelies/pysos\" target=\"_blank\">https://github.com/dagnelies/pysos</a> worked just fine</p>",
      "rawMarkdown": "In any case someone is looking for a very easy way, my submission using https://github.com/dagnelies/pysos worked just fine",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1141343,
      "author_name": "mingpan07",
      "author_url": "",
      "post_date": "01/06/2021 16:37:36",
      "content": "<p>Be careful if you're going to use this - I tried it at the start of the competition a few months ago but found there were some bugs (unfortunately I don't remember what they were). You may still find it useful, just don't blindly trust it to solve all your problems.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1141351,
          "author_name": "rodolphelampe",
          "author_url": "",
          "post_date": "01/06/2021 16:43:02",
          "content": "<p>Thanks. Any alternative ?<br>\nI had some errors because I used numpy types (which are not json serializable), do you remember if it was related to that ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1141822,
          "author_name": "rodolphelampe",
          "author_url": "",
          "post_date": "01/06/2021 23:19:00",
          "content": "<p>In any case someone is looking for a very easy way, my submission using <a href=\"https://github.com/dagnelies/pysos\" target=\"_blank\">https://github.com/dagnelies/pysos</a> worked just fine</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1141391,
      "author_name": "stmandl",
      "author_url": "",
      "post_date": "01/06/2021 17:15:34",
      "content": "<p>I'm using SQLite. I was surprised how fast it is and it let's you maintain large statistics. With some wrapper code it almost feels like an infinite dictionary.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1141402,
          "author_name": "yangxiaoshuai",
          "author_url": "",
          "post_date": "01/06/2021 17:25:20",
          "content": "<p>may I know how to store double nested dictionaries with sqlite? concat two keys as one key?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1141420,
          "author_name": "stmandl",
          "author_url": "",
          "post_date": "01/06/2021 17:38:52",
          "content": "<p>Sure, it's even simpler: just use two columns as key.</p>\n<p>Basically like this:</p>\n<p><code>update user_question_table set key=value where user_id=... and question_id=...</code></p>\n<p>To get fast retrievals it's a good idea to also add a multi-column index on the table.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1141436,
          "author_name": "yangxiaoshuai",
          "author_url": "",
          "post_date": "01/06/2021 17:50:22",
          "content": "<p>thank you 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1141401,
      "author_name": "yangxiaoshuai",
      "author_url": "",
      "post_date": "01/06/2021 17:23:39",
      "content": "<p>I still tend to use sqlite, but double nested dictionaries do not know how to store with sqlite</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1141328": "I spent quite some time trying to figure out how to decrease RAM usage. I didn't like the idea of using SQLite as the syntax is not easy and it's not easy to use dicts for that. I tried the module shelve (https://docs.python.org/3/library/shelve.html) which is exactly what I needed but unfortunately kaggle environnement doesn't have the backend ndbm which makes it impossible to use.\nI found this repo https://github.com/dagnelies/pysos which is quite awesome, it makes the job and doesn't need any backend. It's very fast and let's you use any kind of jsonable object (think dicts and lists).",
    "1141343": "Be careful if you're going to use this - I tried it at the start of the competition a few months ago but found there were some bugs (unfortunately I don't remember what they were). You may still find it useful, just don't blindly trust it to solve all your problems.",
    "1141351": "Thanks. Any alternative ?\nI had some errors because I used numpy types (which are not json serializable), do you remember if it was related to that ?",
    "1141391": "I'm using SQLite. I was surprised how fast it is and it let's you maintain large statistics. With some wrapper code it almost feels like an infinite dictionary.",
    "1141401": "I still tend to use sqlite, but double nested dictionaries do not know how to store with sqlite",
    "1141402": "may I know how to store double nested dictionaries with sqlite? concat two keys as one key?",
    "1141420": "Sure, it's even simpler: just use two columns as key.\n\nBasically like this:\n\n`update user_question_table set key=value where user_id=... and question_id=...`\n\nTo get fast retrievals it's a good idea to also add a multi-column index on the table.",
    "1141436": "thank you 👍",
    "1141822": "In any case someone is looking for a very easy way, my submission using https://github.com/dagnelies/pysos worked just fine"
  },
  "source": "meta"
}