{
  "id": 131794,
  "title": "Effective way to store/load video frames?",
  "url": "/competitions/deepfake-detection-challenge/discussion/131794",
  "author_name": "",
  "post_date": "2020-02-21T18:32:12.087631500Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>what are the effective ways to store/load video frames for model training purposes?</p>",
  "messages": [
    {
      "id": "753124",
      "postDate": "02/21/2020 18:32:12",
      "content": "<p>what are the effective ways to store/load video frames for model training purposes?</p>",
      "rawMarkdown": "what are the effective ways to store/load video frames for model training purposes?",
      "votes": null
    },
    {
      "id": "754200",
      "postDate": "02/23/2020 07:51:43",
      "content": "<p>I have been using sqlitedict (<a href=\"https://github.com/RaRe-Technologies/sqlitedict\">https://github.com/RaRe-Technologies/sqlitedict</a>) it is pretty easy to set up, you install tit via pip easily and it behaves like a dictionary that is persisted on your disk.</p>\n\n<p>For inference (submissions) I just remove that code from my notebook and just generate features without it.</p>\n\n<p>So for example, if you want to create a dataset for training, you can just put features into that dictionary, a filename can be your key, and the value can be your feature. You can get fancy and version your features and add some metadata so you can always keep track of what you have done.</p>\n\n<p>```python\nimport sqlitedict</p>\n\n<p>CACHE = sqlitedict.SqliteDict('deepfake_features_v1.db', autocommit=True)</p>\n\n<h1>somewhere in your dataset code</h1>\n\n<p>.....\nfor filename, target in zip(filenames, targets):\n  if filename in CACHE:\n    features = CACHE[filename]\n  else:\n    features = self.get_features(filename)\n    CACHE[filename] = features\n```</p>",
      "rawMarkdown": "I have been using sqlitedict (https://github.com/RaRe-Technologies/sqlitedict) it is pretty easy to set up, you install tit via pip easily and it behaves like a dictionary that is persisted on your disk.\n\nFor inference (submissions) I just remove that code from my notebook and just generate features without it.\n\nSo for example, if you want to create a dataset for training, you can just put features into that dictionary, a filename can be your key, and the value can be your feature. You can get fancy and version your features and add some metadata so you can always keep track of what you have done.\n\n```python\nimport sqlitedict\n\nCACHE = sqlitedict.SqliteDict('deepfake_features_v1.db', autocommit=True)\n\n# somewhere in your dataset code\n.....\nfor filename, target in zip(filenames, targets):\n  if filename in CACHE:\n    features = CACHE[filename]\n  else:\n    features = self.get_features(filename)\n    CACHE[filename] = features\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 754200,
      "author_name": "weenkus",
      "author_url": "",
      "post_date": "02/23/2020 07:51:43",
      "content": "<p>I have been using sqlitedict (<a href=\"https://github.com/RaRe-Technologies/sqlitedict\">https://github.com/RaRe-Technologies/sqlitedict</a>) it is pretty easy to set up, you install tit via pip easily and it behaves like a dictionary that is persisted on your disk.</p>\n\n<p>For inference (submissions) I just remove that code from my notebook and just generate features without it.</p>\n\n<p>So for example, if you want to create a dataset for training, you can just put features into that dictionary, a filename can be your key, and the value can be your feature. You can get fancy and version your features and add some metadata so you can always keep track of what you have done.</p>\n\n<p>```python\nimport sqlitedict</p>\n\n<p>CACHE = sqlitedict.SqliteDict('deepfake_features_v1.db', autocommit=True)</p>\n\n<h1>somewhere in your dataset code</h1>\n\n<p>.....\nfor filename, target in zip(filenames, targets):\n  if filename in CACHE:\n    features = CACHE[filename]\n  else:\n    features = self.get_features(filename)\n    CACHE[filename] = features\n```</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "753124": "what are the effective ways to store/load video frames for model training purposes?",
    "754200": "I have been using sqlitedict (https://github.com/RaRe-Technologies/sqlitedict) it is pretty easy to set up, you install tit via pip easily and it behaves like a dictionary that is persisted on your disk.\n\nFor inference (submissions) I just remove that code from my notebook and just generate features without it.\n\nSo for example, if you want to create a dataset for training, you can just put features into that dictionary, a filename can be your key, and the value can be your feature. You can get fancy and version your features and add some metadata so you can always keep track of what you have done.\n\n```python\nimport sqlitedict\n\nCACHE = sqlitedict.SqliteDict('deepfake_features_v1.db', autocommit=True)\n\n# somewhere in your dataset code\n.....\nfor filename, target in zip(filenames, targets):\n  if filename in CACHE:\n    features = CACHE[filename]\n  else:\n    features = self.get_features(filename)\n    CACHE[filename] = features\n```"
  },
  "source": "meta"
}