{
  "id": 211638,
  "title": "Simple trick if you need more RAM",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/211638",
  "author_name": "",
  "post_date": "2021-01-15T22:47:49.858896700Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The RAM limitations in the notebooks mean that you can load an image to memory and construct an int8 mask one tile at a time, something like the code below. </p>\n<pre><code>pred_mask = np.zeros(dataset.shape, dtype=np.uint8)\n\nfor (x1, x2, y1, y2) in tqdm(slices, desc=img_path.stem):\n    image = dataset.read([1, 2, 3], window=Window.from_slices((x1, x2), (y1, y2)))\n    image = np.moveaxis(image, 0, -1)\n    image = trfm(image)\n\n    predictions = 0\n\n    for model in models:\n        model.to(device)\n        model.eval()\n\n        with torch.no_grad():\n            img = image.to(device)[None]\n            predictions += model(img)\n\n            # TTA etc...\n\n    predictions /= len(models)\n    predictions = torch.squeeze(predictions).sigmoid().cpu().numpy()\n    predictions = cv2.resize(predictions, (tile_size, tile_size))\n    pred_mask[x1:x2, y1:y2] = predictions &gt; 0.5\n\nreturn pred_mask\n</code></pre>\n<p>But if you want to use the full prediction array (float), you might run into out-of-memory errors when you exceed the 13GB RAM limit.</p>\n<p>However, we also have 16GB of VRAM at our disposal, so we can keep a full resolution float16 array of our predictions/probabilities on the GPU, something like this:</p>\n<pre><code>pred_probs = torch.zeros(dataset.shape, dtype=torch.float16).to(device)\n\nfor (x1, x2, y1, y2) in tqdm(slices, desc=img_path.stem):\n    image = dataset.read([1, 2, 3], window=Window.from_slices((x1, x2), (y1, y2)))\n    image = np.moveaxis(image, 0, -1)\n    image = trfm(image)\n\n    predictions = 0\n\n    for model in models:\n        model.to(device)\n        model.eval()\n\n        with torch.no_grad():\n            img = image.to(device)[None]\n            predictions += model(img)\n\n            # TTA etc...\n\n    predictions /= len(models)\n    predictions = resize(predictions.sigmoid(), (tile_size, tile_size))  # torchvision.transforms\n    pred_probs[x1:x2, y1:y2] += torch.squeeze(predictions)\n\n# Do something with probs\npred_probs = magic_postprocess(pred_probs)\n\nbool_mask = pred_probs &gt; 0.5\nreturn bool_mask.cpu().numpy()\n</code></pre>\n<p>This might come in handy for postprocessing where you need to take into account multiple tiles (e.g. overlap smoothing etc.)</p>",
  "messages": [
    {
      "id": "1154796",
      "postDate": "01/15/2021 22:47:49",
      "content": "<p>The RAM limitations in the notebooks mean that you can load an image to memory and construct an int8 mask one tile at a time, something like the code below. </p>\n<pre><code>pred_mask = np.zeros(dataset.shape, dtype=np.uint8)\n\nfor (x1, x2, y1, y2) in tqdm(slices, desc=img_path.stem):\n    image = dataset.read([1, 2, 3], window=Window.from_slices((x1, x2), (y1, y2)))\n    image = np.moveaxis(image, 0, -1)\n    image = trfm(image)\n\n    predictions = 0\n\n    for model in models:\n        model.to(device)\n        model.eval()\n\n        with torch.no_grad():\n            img = image.to(device)[None]\n            predictions += model(img)\n\n            # TTA etc...\n\n    predictions /= len(models)\n    predictions = torch.squeeze(predictions).sigmoid().cpu().numpy()\n    predictions = cv2.resize(predictions, (tile_size, tile_size))\n    pred_mask[x1:x2, y1:y2] = predictions &gt; 0.5\n\nreturn pred_mask\n</code></pre>\n<p>But if you want to use the full prediction array (float), you might run into out-of-memory errors when you exceed the 13GB RAM limit.</p>\n<p>However, we also have 16GB of VRAM at our disposal, so we can keep a full resolution float16 array of our predictions/probabilities on the GPU, something like this:</p>\n<pre><code>pred_probs = torch.zeros(dataset.shape, dtype=torch.float16).to(device)\n\nfor (x1, x2, y1, y2) in tqdm(slices, desc=img_path.stem):\n    image = dataset.read([1, 2, 3], window=Window.from_slices((x1, x2), (y1, y2)))\n    image = np.moveaxis(image, 0, -1)\n    image = trfm(image)\n\n    predictions = 0\n\n    for model in models:\n        model.to(device)\n        model.eval()\n\n        with torch.no_grad():\n            img = image.to(device)[None]\n            predictions += model(img)\n\n            # TTA etc...\n\n    predictions /= len(models)\n    predictions = resize(predictions.sigmoid(), (tile_size, tile_size))  # torchvision.transforms\n    pred_probs[x1:x2, y1:y2] += torch.squeeze(predictions)\n\n# Do something with probs\npred_probs = magic_postprocess(pred_probs)\n\nbool_mask = pred_probs &gt; 0.5\nreturn bool_mask.cpu().numpy()\n</code></pre>\n<p>This might come in handy for postprocessing where you need to take into account multiple tiles (e.g. overlap smoothing etc.)</p>",
      "rawMarkdown": "The RAM limitations in the notebooks mean that you can load an image to memory and construct an int8 mask one tile at a time, something like the code below. \n\n```\npred_mask = np.zeros(dataset.shape, dtype=np.uint8)\n\nfor (x1, x2, y1, y2) in tqdm(slices, desc=img_path.stem):\n    image = dataset.read([1, 2, 3], window=Window.from_slices((x1, x2), (y1, y2)))\n    image = np.moveaxis(image, 0, -1)\n    image = trfm(image)\n\n    predictions = 0\n\n    for model in models:\n        model.to(device)\n        model.eval()\n\n        with torch.no_grad():\n            img = image.to(device)[None]\n            predictions += model(img)\n\n            # TTA etc...\n\n    predictions /= len(models)\n    predictions = torch.squeeze(predictions).sigmoid().cpu().numpy()\n    predictions = cv2.resize(predictions, (tile_size, tile_size))\n    pred_mask[x1:x2, y1:y2] = predictions > 0.5\n\nreturn pred_mask\n```\nBut if you want to use the full prediction array (float), you might run into out-of-memory errors when you exceed the 13GB RAM limit.\n\nHowever, we also have 16GB of VRAM at our disposal, so we can keep a full resolution float16 array of our predictions/probabilities on the GPU, something like this:\n\n```\npred_probs = torch.zeros(dataset.shape, dtype=torch.float16).to(device)\n\nfor (x1, x2, y1, y2) in tqdm(slices, desc=img_path.stem):\n    image = dataset.read([1, 2, 3], window=Window.from_slices((x1, x2), (y1, y2)))\n    image = np.moveaxis(image, 0, -1)\n    image = trfm(image)\n\n    predictions = 0\n\n    for model in models:\n        model.to(device)\n        model.eval()\n\n        with torch.no_grad():\n            img = image.to(device)[None]\n            predictions += model(img)\n            \n            # TTA etc...\n\n    predictions /= len(models)\n    predictions = resize(predictions.sigmoid(), (tile_size, tile_size))  # torchvision.transforms\n    pred_probs[x1:x2, y1:y2] += torch.squeeze(predictions)\n\n# Do something with probs\npred_probs = magic_postprocess(pred_probs)\n\nbool_mask = pred_probs > 0.5\nreturn bool_mask.cpu().numpy()\n```\nThis might come in handy for postprocessing where you need to take into account multiple tiles (e.g. overlap smoothing etc.)",
      "votes": null
    },
    {
      "id": "1155290",
      "postDate": "01/16/2021 10:55:35",
      "content": "<p>Nice one, I shall give it a try</p>",
      "rawMarkdown": "Nice one, I shall give it a try",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1155290,
      "author_name": "nageshsingh",
      "author_url": "",
      "post_date": "01/16/2021 10:55:35",
      "content": "<p>Nice one, I shall give it a try</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1154796": "The RAM limitations in the notebooks mean that you can load an image to memory and construct an int8 mask one tile at a time, something like the code below. \n\n```\npred_mask = np.zeros(dataset.shape, dtype=np.uint8)\n\nfor (x1, x2, y1, y2) in tqdm(slices, desc=img_path.stem):\n    image = dataset.read([1, 2, 3], window=Window.from_slices((x1, x2), (y1, y2)))\n    image = np.moveaxis(image, 0, -1)\n    image = trfm(image)\n\n    predictions = 0\n\n    for model in models:\n        model.to(device)\n        model.eval()\n\n        with torch.no_grad():\n            img = image.to(device)[None]\n            predictions += model(img)\n\n            # TTA etc...\n\n    predictions /= len(models)\n    predictions = torch.squeeze(predictions).sigmoid().cpu().numpy()\n    predictions = cv2.resize(predictions, (tile_size, tile_size))\n    pred_mask[x1:x2, y1:y2] = predictions > 0.5\n\nreturn pred_mask\n```\nBut if you want to use the full prediction array (float), you might run into out-of-memory errors when you exceed the 13GB RAM limit.\n\nHowever, we also have 16GB of VRAM at our disposal, so we can keep a full resolution float16 array of our predictions/probabilities on the GPU, something like this:\n\n```\npred_probs = torch.zeros(dataset.shape, dtype=torch.float16).to(device)\n\nfor (x1, x2, y1, y2) in tqdm(slices, desc=img_path.stem):\n    image = dataset.read([1, 2, 3], window=Window.from_slices((x1, x2), (y1, y2)))\n    image = np.moveaxis(image, 0, -1)\n    image = trfm(image)\n\n    predictions = 0\n\n    for model in models:\n        model.to(device)\n        model.eval()\n\n        with torch.no_grad():\n            img = image.to(device)[None]\n            predictions += model(img)\n            \n            # TTA etc...\n\n    predictions /= len(models)\n    predictions = resize(predictions.sigmoid(), (tile_size, tile_size))  # torchvision.transforms\n    pred_probs[x1:x2, y1:y2] += torch.squeeze(predictions)\n\n# Do something with probs\npred_probs = magic_postprocess(pred_probs)\n\nbool_mask = pred_probs > 0.5\nreturn bool_mask.cpu().numpy()\n```\nThis might come in handy for postprocessing where you need to take into account multiple tiles (e.g. overlap smoothing etc.)",
    "1155290": "Nice one, I shall give it a try"
  },
  "source": "meta"
}