{
  "id": 436710,
  "title": "npz taking a long time to load?",
  "url": "/competitions/predict-ai-model-runtime/discussion/436710",
  "author_name": "",
  "post_date": "2023-09-03T17:28:51.295282800Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I am simply iterating through all the files in the dataset and loading them using dict(np.load()) but am noticing that the time taken to load these files is huge - the script has been running for 16 hours and it's still going.</p>\n<p>(<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16562613%2F99683f4b1c3054889add6baa2bd0cbf2%2FScreenshot%20from%202023-09-03%2010-25-51.png?generation=1693762032360137&amp;alt=media\" target=\"_blank\">https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16562613%2F99683f4b1c3054889add6baa2bd0cbf2%2FScreenshot%20from%202023-09-03%2010-25-51.png?generation=1693762032360137&amp;alt=media</a>)</p>\n<p>I looked into which exact files are taking the longest time to load and I think it's the ones in /layout. Are other folks seeing the same thing? Wondering how would we train on them if it takes so long just to load them into memory. </p>\n<p>Clarification: I am converting the loaded npz files into dict as a part of this step as well.</p>",
  "messages": [
    {
      "id": "2422074",
      "postDate": "09/03/2023 17:28:51",
      "content": "<p>I am simply iterating through all the files in the dataset and loading them using dict(np.load()) but am noticing that the time taken to load these files is huge - the script has been running for 16 hours and it's still going.</p>\n<p>(<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16562613%2F99683f4b1c3054889add6baa2bd0cbf2%2FScreenshot%20from%202023-09-03%2010-25-51.png?generation=1693762032360137&amp;alt=media\" target=\"_blank\">https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16562613%2F99683f4b1c3054889add6baa2bd0cbf2%2FScreenshot%20from%202023-09-03%2010-25-51.png?generation=1693762032360137&amp;alt=media</a>)</p>\n<p>I looked into which exact files are taking the longest time to load and I think it's the ones in /layout. Are other folks seeing the same thing? Wondering how would we train on them if it takes so long just to load them into memory. </p>\n<p>Clarification: I am converting the loaded npz files into dict as a part of this step as well.</p>",
      "rawMarkdown": "I am simply iterating through all the files in the dataset and loading them using dict(np.load(<file_path>)) but am noticing that the time taken to load these files is huge - the script has been running for 16 hours and it's still going.\n\n(https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16562613%2F99683f4b1c3054889add6baa2bd0cbf2%2FScreenshot%20from%202023-09-03%2010-25-51.png?generation=1693762032360137&alt=media)\n\nI looked into which exact files are taking the longest time to load and I think it's the ones in /layout. Are other folks seeing the same thing? Wondering how would we train on them if it takes so long just to load them into memory. \n\nClarification: I am converting the loaded npz files into dict as a part of this step as well.",
      "votes": null
    },
    {
      "id": "2426077",
      "postDate": "09/06/2023 11:46:12",
      "content": "<p>Hello,</p>\n<p>You can have a decent speedup by only reading the <code>key</code> that you're interested in. In NPZ file format, the heavy reading of the data is actually performed when you're accessing the value from specified key. If you're not accessing a certain key, then the value of that key will not be read, thus saving your time.</p>\n<p>For example, if I only interested in reading <code>node_feat</code> and <code>node_opcode</code>, then I will simply do:</p>\n<pre><code>npz = np.load()\nnode_feat = npz[]\nnode_opcode = npz[]\n\n</code></pre>",
      "rawMarkdown": "Hello,\n\nYou can have a decent speedup by only reading the `key` that you're interested in. In NPZ file format, the heavy reading of the data is actually performed when you're accessing the value from specified key. If you're not accessing a certain key, then the value of that key will not be read, thus saving your time.\n\nFor example, if I only interested in reading `node_feat` and `node_opcode`, then I will simply do:\n\n```python\nnpz = np.load(\"npz/layout/xla/random/train/<pick_1>.npz\")\nnode_feat = npz['node_feat']\nnode_opcode = npz['node_opcode']\n# faster because any other keys are not being read\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2426077,
      "author_name": "thariqnugrohotomo",
      "author_url": "",
      "post_date": "09/06/2023 11:46:12",
      "content": "<p>Hello,</p>\n<p>You can have a decent speedup by only reading the <code>key</code> that you're interested in. In NPZ file format, the heavy reading of the data is actually performed when you're accessing the value from specified key. If you're not accessing a certain key, then the value of that key will not be read, thus saving your time.</p>\n<p>For example, if I only interested in reading <code>node_feat</code> and <code>node_opcode</code>, then I will simply do:</p>\n<pre><code>npz = np.load()\nnode_feat = npz[]\nnode_opcode = npz[]\n\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2422074": "I am simply iterating through all the files in the dataset and loading them using dict(np.load(<file_path>)) but am noticing that the time taken to load these files is huge - the script has been running for 16 hours and it's still going.\n\n(https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16562613%2F99683f4b1c3054889add6baa2bd0cbf2%2FScreenshot%20from%202023-09-03%2010-25-51.png?generation=1693762032360137&alt=media)\n\nI looked into which exact files are taking the longest time to load and I think it's the ones in /layout. Are other folks seeing the same thing? Wondering how would we train on them if it takes so long just to load them into memory. \n\nClarification: I am converting the loaded npz files into dict as a part of this step as well.",
    "2426077": "Hello,\n\nYou can have a decent speedup by only reading the `key` that you're interested in. In NPZ file format, the heavy reading of the data is actually performed when you're accessing the value from specified key. If you're not accessing a certain key, then the value of that key will not be read, thus saving your time.\n\nFor example, if I only interested in reading `node_feat` and `node_opcode`, then I will simply do:\n\n```python\nnpz = np.load(\"npz/layout/xla/random/train/<pick_1>.npz\")\nnode_feat = npz['node_feat']\nnode_opcode = npz['node_opcode']\n# faster because any other keys are not being read\n```"
  },
  "source": "meta"
}