{
  "id": 242358,
  "title": "Memory leak when reading .npy in a loop",
  "url": "/competitions/seti-breakthrough-listen/discussion/242358",
  "author_name": "Adriano Passos",
  "post_date": "2021-05-28T15:20:11.732000",
  "votes": 0,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I am crashing my collab session on inference due to looping my data loader. For some reason, my loader was piling up the np objects and was never getting rid of it.</p>\n<p>Even with gc it still crashes. Somehow it works during training inside Pytorch Lightning. Any tips?</p>\n<pre><code>import gc\nclass alienDataset(Dataset):\n  def __init__(self, base_dir, ids):\n    super().__init__()\n    self.ids = ids\n    self.base_dir = base_dir\n\n  def __len__(self):\n    return len(self.ids)\n\n  def __getitem__(self, idx):\n    id = self.ids[idx]\n    path = self.base_dir + id[0] + '/' + id + '.npy'\n    x_np = np.load(path)\n    x = torch.tensor(x_np).float()\n    x = transforms.CenterCrop((255,255))(x)\n    del x_np\n    gc.collect()\n\n    return x\n</code></pre>\n<p>when I run a simple inference code my ram explodes</p>\n<pre><code>preds = []\nfor xb in tqdm(train_loader):\n  preds += [model(xb)]\n</code></pre>",
  "messages": [
    {
      "id": 1327126,
      "postDate": "2021-05-29T00:25:45.300Z",
      "content": "<p>can you use memmap?<br>\nmapping all file before dataset with np.load(path, mmap_mode = \"r\")</p>",
      "rawMarkdown": "can you use memmap?\nmapping all file before dataset with np.load(path, mmap_mode = \"r\")",
      "replies": [
        {
          "id": 1327131,
          "postDate": "2021-05-29T00:40:09.027Z",
          "content": "<p>Tnx for the reply. I fixed the issue already. In the end it had nothing to do with numpy… I simply forgot to call call 'model.eval() / torch.no_grad()' before the inference loop.</p>",
          "rawMarkdown": "Tnx for the reply. I fixed the issue already. In the end it had nothing to do with numpy... I simply forgot to call call 'model.eval() / torch.no_grad()' before the inference loop."
        }
      ]
    },
    {
      "id": 1326655,
      "postDate": "2021-05-28T15:44:07.080Z",
      "content": "<p>Apparently, the leak is in my model, not in numpy. When I run</p>\n<pre><code>for x in tqdm(train_loader):\n  pass\n</code></pre>\n<p>there is no leak.</p>",
      "rawMarkdown": "Apparently, the leak is in my model, not in numpy. When I run\n```\nfor x in tqdm(train_loader):\n  pass\n```\nthere is no leak.",
      "replies": [
        {
          "id": 1326662,
          "postDate": "2021-05-28T15:46:07.190Z",
          "content": "<p>Ok, I feel so dumb now… I forgot to call model.eval() / torch.no_grad() 🙈</p>",
          "rawMarkdown": "Ok, I feel so dumb now... I forgot to call model.eval() / torch.no_grad() 🙈"
        }
      ]
    },
    {
      "id": 1326650,
      "postDate": "2021-05-28T15:39:00.510Z",
      "content": "<p>It also worth mentioning that the first epoch on Pytorch Lightning takes significantly more time than the other ones, Maybe python is cashing the dataset somehow? but the trainer manages to use that properly, but when I loop over it manually it goes mayhem?</p>",
      "rawMarkdown": "It also worth mentioning that the first epoch on Pytorch Lightning takes significantly more time than the other ones, Maybe python is cashing the dataset somehow? but the trainer manages to use that properly, but when I loop over it manually it goes mayhem?"
    },
    {
      "id": 1326634,
      "postDate": "2021-05-28T15:27:22.473Z",
      "content": "<p>Try at the end of alienDataset</p>\n<p>From:<br>\n<strong>del x_np\ngc.collect()</strong></p>\n<p>To:<br>\n<strong>del x_np\nx_np.close()\ngc.collect()</strong></p>",
      "rawMarkdown": "Try at the end of alienDataset\n\nFrom:\n**del x_np\ngc.collect()**\n\nTo:\n**del x_np\nx_np.close()\ngc.collect()**",
      "replies": [
        {
          "id": 1326637,
          "postDate": "2021-05-28T15:29:22.297Z",
          "content": "<p>Tnx for the quick response. I get the following</p>\n<blockquote>\n  <p>UnboundLocalError: local variable 'x_np' referenced before assignment</p>\n</blockquote>\n<p>and if I reverse the order from del and close I get the following: </p>\n<blockquote>\n  <p>AttributeError: 'numpy.ndarray' object has no attribute 'close'</p>\n</blockquote>",
          "rawMarkdown": "Tnx for the quick response. I get the following\n\n> UnboundLocalError: local variable 'x_np' referenced before assignment\n\nand if I reverse the order from del and close I get the following: \n\n> AttributeError: 'numpy.ndarray' object has no attribute 'close'"
        },
        {
          "id": 1326643,
          "postDate": "2021-05-28T15:33:13.287Z",
          "content": "<p>Sorry missed the .f at the end. Please try this.</p>\n<p>del x_np.f <br>\nx_np.close() </p>\n<p>gc.collect()</p>",
          "rawMarkdown": "Sorry missed the .f at the end. Please try this.\n\n\ndel x_np.f \nx_np.close() \n\ngc.collect()"
        },
        {
          "id": 1326645,
          "postDate": "2021-05-28T15:34:12.853Z",
          "content": "<blockquote>\n  <p>AttributeError: 'numpy.ndarray' object has no attribute 'f'</p>\n</blockquote>",
          "rawMarkdown": ">AttributeError: 'numpy.ndarray' object has no attribute 'f'"
        }
      ]
    },
    {
      "id": 1326624,
      "postDate": "2021-05-28T15:20:11.733Z",
      "content": "<p>I am crashing my collab session on inference due to looping my data loader. For some reason, my loader was piling up the np objects and was never getting rid of it.</p>\n<p>Even with gc it still crashes. Somehow it works during training inside Pytorch Lightning. Any tips?</p>\n<pre><code>import gc\nclass alienDataset(Dataset):\n  def __init__(self, base_dir, ids):\n    super().__init__()\n    self.ids = ids\n    self.base_dir = base_dir\n\n  def __len__(self):\n    return len(self.ids)\n\n  def __getitem__(self, idx):\n    id = self.ids[idx]\n    path = self.base_dir + id[0] + '/' + id + '.npy'\n    x_np = np.load(path)\n    x = torch.tensor(x_np).float()\n    x = transforms.CenterCrop((255,255))(x)\n    del x_np\n    gc.collect()\n\n    return x\n</code></pre>\n<p>when I run a simple inference code my ram explodes</p>\n<pre><code>preds = []\nfor xb in tqdm(train_loader):\n  preds += [model(xb)]\n</code></pre>",
      "rawMarkdown": "I am crashing my collab session on inference due to looping my data loader. For some reason, my loader was piling up the np objects and was never getting rid of it.\n\nEven with gc it still crashes. Somehow it works during training inside Pytorch Lightning. Any tips?\n\n```\nimport gc\nclass alienDataset(Dataset):\n  def __init__(self, base_dir, ids):\n    super().__init__()\n    self.ids = ids\n    self.base_dir = base_dir\n\n  def __len__(self):\n    return len(self.ids)\n\n  def __getitem__(self, idx):\n    id = self.ids[idx]\n    path = self.base_dir + id[0] + '/' + id + '.npy'\n    x_np = np.load(path)\n    x = torch.tensor(x_np).float()\n    x = transforms.CenterCrop((255,255))(x)\n    del x_np\n    gc.collect()\n    \n    return x\n```\n\nwhen I run a simple inference code my ram explodes\n\n```\npreds = []\nfor xb in tqdm(train_loader):\n  preds += [model(xb)]\n```\n"
    }
  ],
  "comments": [
    {
      "id": 1327126,
      "author_name": "assign",
      "author_url": "",
      "post_date": "2021-05-29T00:25:45.300000",
      "content": "<p>can you use memmap?<br>\nmapping all file before dataset with np.load(path, mmap_mode = \"r\")</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1327131,
          "author_name": "Adriano Passos",
          "author_url": "",
          "post_date": "2021-05-29T00:40:09.027000",
          "content": "<p>Tnx for the reply. I fixed the issue already. In the end it had nothing to do with numpy… I simply forgot to call call 'model.eval() / torch.no_grad()' before the inference loop.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1326655,
      "author_name": "Adriano Passos",
      "author_url": "",
      "post_date": "2021-05-28T15:44:07.080000",
      "content": "<p>Apparently, the leak is in my model, not in numpy. When I run</p>\n<pre><code>for x in tqdm(train_loader):\n  pass\n</code></pre>\n<p>there is no leak.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1326662,
          "author_name": "Adriano Passos",
          "author_url": "",
          "post_date": "2021-05-28T15:46:07.190000",
          "content": "<p>Ok, I feel so dumb now… I forgot to call model.eval() / torch.no_grad() 🙈</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1326650,
      "author_name": "Adriano Passos",
      "author_url": "",
      "post_date": "2021-05-28T15:39:00.510000",
      "content": "<p>It also worth mentioning that the first epoch on Pytorch Lightning takes significantly more time than the other ones, Maybe python is cashing the dataset somehow? but the trainer manages to use that properly, but when I loop over it manually it goes mayhem?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1326634,
      "author_name": "Fayzur",
      "author_url": "",
      "post_date": "2021-05-28T15:27:22.473000",
      "content": "<p>Try at the end of alienDataset</p>\n<p>From:<br>\n<strong>del x_np\ngc.collect()</strong></p>\n<p>To:<br>\n<strong>del x_np\nx_np.close()\ngc.collect()</strong></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1326637,
          "author_name": "Adriano Passos",
          "author_url": "",
          "post_date": "2021-05-28T15:29:22.297000",
          "content": "<p>Tnx for the quick response. I get the following</p>\n<blockquote>\n  <p>UnboundLocalError: local variable 'x_np' referenced before assignment</p>\n</blockquote>\n<p>and if I reverse the order from del and close I get the following: </p>\n<blockquote>\n  <p>AttributeError: 'numpy.ndarray' object has no attribute 'close'</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1326643,
          "author_name": "Fayzur",
          "author_url": "",
          "post_date": "2021-05-28T15:33:13.287000",
          "content": "<p>Sorry missed the .f at the end. Please try this.</p>\n<p>del x_np.f <br>\nx_np.close() </p>\n<p>gc.collect()</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1326645,
          "author_name": "Adriano Passos",
          "author_url": "",
          "post_date": "2021-05-28T15:34:12.853000",
          "content": "<blockquote>\n  <p>AttributeError: 'numpy.ndarray' object has no attribute 'f'</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1327126": "can you use memmap?\nmapping all file before dataset with np.load(path, mmap_mode = \"r\")",
    "1326655": "Apparently, the leak is in my model, not in numpy. When I run\n```\nfor x in tqdm(train_loader):\n  pass\n```\nthere is no leak.",
    "1326650": "It also worth mentioning that the first epoch on Pytorch Lightning takes significantly more time than the other ones, Maybe python is cashing the dataset somehow? but the trainer manages to use that properly, but when I loop over it manually it goes mayhem?",
    "1326634": "Try at the end of alienDataset\n\nFrom:\n**del x_np\ngc.collect()**\n\nTo:\n**del x_np\nx_np.close()\ngc.collect()**",
    "1326624": "I am crashing my collab session on inference due to looping my data loader. For some reason, my loader was piling up the np objects and was never getting rid of it.\n\nEven with gc it still crashes. Somehow it works during training inside Pytorch Lightning. Any tips?\n\n```\nimport gc\nclass alienDataset(Dataset):\n  def __init__(self, base_dir, ids):\n    super().__init__()\n    self.ids = ids\n    self.base_dir = base_dir\n\n  def __len__(self):\n    return len(self.ids)\n\n  def __getitem__(self, idx):\n    id = self.ids[idx]\n    path = self.base_dir + id[0] + '/' + id + '.npy'\n    x_np = np.load(path)\n    x = torch.tensor(x_np).float()\n    x = transforms.CenterCrop((255,255))(x)\n    del x_np\n    gc.collect()\n    \n    return x\n```\n\nwhen I run a simple inference code my ram explodes\n\n```\npreds = []\nfor xb in tqdm(train_loader):\n  preds += [model(xb)]\n```\n"
  }
}