{
  "id": 405515,
  "title": "RAM out of memory ",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/405515",
  "author_name": "",
  "post_date": "2023-04-28T04:52:04.966371700Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have taken middle 7 slices out of all the 65 slices and cut off the image 3d volume into 224x224 image slices but whenever I go for image 2 with seresnext model my RAM memory is full can someone provide any advice what shall I do </p>",
  "messages": [
    {
      "id": "2237927",
      "postDate": "04/28/2023 04:52:04",
      "content": "<p>I have taken middle 7 slices out of all the 65 slices and cut off the image 3d volume into 224x224 image slices but whenever I go for image 2 with seresnext model my RAM memory is full can someone provide any advice what shall I do </p>",
      "rawMarkdown": "I have taken middle 7 slices out of all the 65 slices and cut off the image 3d volume into 224x224 image slices but whenever I go for image 2 with seresnext model my RAM memory is full can someone provide any advice what shall I do",
      "votes": null
    },
    {
      "id": "2238371",
      "postDate": "04/28/2023 13:14:49",
      "content": "<p>There are several tips I'm using in data loading for this competition:</p>\n<p><strong>- Preallocate memory instead of concatenation</strong><br>\n<code>np.concatenate</code> or any other method to retrieve a single array from a list of many small arrays uses some overhead memory. The best thing to deal with it is to pre-allocate the array instead:</p>\n<pre><code> ():\n    tile_idxs = get_tile_idxs()\n    data = np.array(((tile_idxs), Config.TILE_SIZE, Config.TILE_SIZE, Config.N_SLICES), dtype=np.uint8)\n     i, slice_path  (slice_paths):\n        slice_data = load_slice(slice_path)\n        data[..., i] = split_slice(slice_data, tile_idxs)\n     data\n\n ():\n    tile_idxs = get_tile_idxs()\n    tiles = []\n     slice_path  slice_paths:\n        slice_data = load_slice(slice_path)\n        tiles.append(split_slice(slice_data, tile_idxs))\n    data = np.concatenate(tiles, axis=))  \n     data\n</code></pre>\n<p><strong>- Assign proper <code>dtype</code> for the volume <code>np.array</code> to reduce memory usage</strong><br>\nThe difference between 16-bit values used in raw tiff files and 8-bit values shouldn't affect the model's performance too much.<br>\nSo, you could store the volume data as <code>np.uint8</code> in the <code>[0-255]</code> range and convert it to <code>[0-1]</code> float in the <code>__getitem__</code> method of your <code>Dataset</code> class to reduce the memory usage constantly consumed by volume data.</p>\n<p><strong>- Don't store tiles outside the mask</strong><br>\nYou could incorporate some logic to skip the tiles outside the <code>mask.png</code> if you didn't implement it yet. It will save you a serious amount of memory since a big part of the images is filled with zeros and doesn't represent anything useful in this competition.</p>\n<hr>\n<p>Other advice could help someone:</p>\n<p><strong>- Use a smaller number of slices</strong><br>\nFor those guys who decided to use all 65 slices, it would take you too much memory, and it's not possible to fit the whole dataset into your RAM, specifically on kaggle. On my setup, the whole dataset with all the tips above but with all 65 slices required at least 22 GB of RAM.</p>\n<p><strong>- Resize the images</strong><br>\nThere are plenty of public notebooks that use that trick in different ways. But in general, you can resize the initial images to reduce memory usage.</p>",
      "rawMarkdown": "There are several tips I'm using in data loading for this competition:\n\n**- Preallocate memory instead of concatenation**\n`np.concatenate` or any other method to retrieve a single array from a list of many small arrays uses some overhead memory. The best thing to deal with it is to pre-allocate the array instead:\n```python\ndef load_tiles_efficient_way():\n    tile_idxs = get_tile_idxs()\n    data = np.array((len(tile_idxs), Config.TILE_SIZE, Config.TILE_SIZE, Config.N_SLICES), dtype=np.uint8)\n    for i, slice_path in enumerate(slice_paths):\n        slice_data = load_slice(slice_path)\n        data[..., i] = split_slice(slice_data, tile_idxs)\n    return data\n\ndef load_tiles_bad_way():\n    tile_idxs = get_tile_idxs()\n    tiles = []\n    for slice_path in slice_paths:\n        slice_data = load_slice(slice_path)\n        tiles.append(split_slice(slice_data, tile_idxs))\n    data = np.concatenate(tiles, axis=0))  # Uses more memory to hold both `tiles` list and `data` array\n    return data\n```\n\n**- Assign proper `dtype` for the volume `np.array` to reduce memory usage**\nThe difference between 16-bit values used in raw tiff files and 8-bit values shouldn't affect the model's performance too much.\nSo, you could store the volume data as `np.uint8` in the `[0-255]` range and convert it to `[0-1]` float in the `__getitem__` method of your `Dataset` class to reduce the memory usage constantly consumed by volume data.\n\n**- Don't store tiles outside the mask**\nYou could incorporate some logic to skip the tiles outside the `mask.png` if you didn't implement it yet. It will save you a serious amount of memory since a big part of the images is filled with zeros and doesn't represent anything useful in this competition.\n\n\n---\n\nOther advice could help someone:\n\n**- Use a smaller number of slices**\nFor those guys who decided to use all 65 slices, it would take you too much memory, and it's not possible to fit the whole dataset into your RAM, specifically on kaggle. On my setup, the whole dataset with all the tips above but with all 65 slices required at least 22 GB of RAM.\n\n**- Resize the images**\nThere are plenty of public notebooks that use that trick in different ways. But in general, you can resize the initial images to reduce memory usage.",
      "votes": null
    },
    {
      "id": "2239768",
      "postDate": "04/29/2023 20:41:17",
      "content": "<p>Is that a problem with the RAM or vRAM (attached to GPU)? </p>\n<p><a href=\"https://www.kaggle.com/r1chardson\" target=\"_blank\">@r1chardson</a> gives a lot of great advice already. I would add: try a smaller backbone and/or smaller batch size if it is a vRAM issue and try gradient accumulation to have a bigger \"virtual\" batch size. </p>",
      "rawMarkdown": "Is that a problem with the RAM or vRAM (attached to GPU)? \n\n@r1chardson gives a lot of great advice already. I would add: try a smaller backbone and/or smaller batch size if it is a vRAM issue and try gradient accumulation to have a bigger \"virtual\" batch size.",
      "votes": null
    },
    {
      "id": "2240085",
      "postDate": "04/30/2023 08:02:42",
      "content": "<p>In my experiments, the best way to prevent RAM OOM errors is to split every big tile into smaller ones and save them on a disk. After that, make inference using the small crops and aggregate the result. </p>",
      "rawMarkdown": "In my experiments, the best way to prevent RAM OOM errors is to split every big tile into smaller ones and save them on a disk. After that, make inference using the small crops and aggregate the result.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2238371,
      "author_name": "r1chardson",
      "author_url": "",
      "post_date": "04/28/2023 13:14:49",
      "content": "<p>There are several tips I'm using in data loading for this competition:</p>\n<p><strong>- Preallocate memory instead of concatenation</strong><br>\n<code>np.concatenate</code> or any other method to retrieve a single array from a list of many small arrays uses some overhead memory. The best thing to deal with it is to pre-allocate the array instead:</p>\n<pre><code> ():\n    tile_idxs = get_tile_idxs()\n    data = np.array(((tile_idxs), Config.TILE_SIZE, Config.TILE_SIZE, Config.N_SLICES), dtype=np.uint8)\n     i, slice_path  (slice_paths):\n        slice_data = load_slice(slice_path)\n        data[..., i] = split_slice(slice_data, tile_idxs)\n     data\n\n ():\n    tile_idxs = get_tile_idxs()\n    tiles = []\n     slice_path  slice_paths:\n        slice_data = load_slice(slice_path)\n        tiles.append(split_slice(slice_data, tile_idxs))\n    data = np.concatenate(tiles, axis=))  \n     data\n</code></pre>\n<p><strong>- Assign proper <code>dtype</code> for the volume <code>np.array</code> to reduce memory usage</strong><br>\nThe difference between 16-bit values used in raw tiff files and 8-bit values shouldn't affect the model's performance too much.<br>\nSo, you could store the volume data as <code>np.uint8</code> in the <code>[0-255]</code> range and convert it to <code>[0-1]</code> float in the <code>__getitem__</code> method of your <code>Dataset</code> class to reduce the memory usage constantly consumed by volume data.</p>\n<p><strong>- Don't store tiles outside the mask</strong><br>\nYou could incorporate some logic to skip the tiles outside the <code>mask.png</code> if you didn't implement it yet. It will save you a serious amount of memory since a big part of the images is filled with zeros and doesn't represent anything useful in this competition.</p>\n<hr>\n<p>Other advice could help someone:</p>\n<p><strong>- Use a smaller number of slices</strong><br>\nFor those guys who decided to use all 65 slices, it would take you too much memory, and it's not possible to fit the whole dataset into your RAM, specifically on kaggle. On my setup, the whole dataset with all the tips above but with all 65 slices required at least 22 GB of RAM.</p>\n<p><strong>- Resize the images</strong><br>\nThere are plenty of public notebooks that use that trick in different ways. But in general, you can resize the initial images to reduce memory usage.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2239768,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "04/29/2023 20:41:17",
      "content": "<p>Is that a problem with the RAM or vRAM (attached to GPU)? </p>\n<p><a href=\"https://www.kaggle.com/r1chardson\" target=\"_blank\">@r1chardson</a> gives a lot of great advice already. I would add: try a smaller backbone and/or smaller batch size if it is a vRAM issue and try gradient accumulation to have a bigger \"virtual\" batch size. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2240085,
      "author_name": "igorkrashenyi",
      "author_url": "",
      "post_date": "04/30/2023 08:02:42",
      "content": "<p>In my experiments, the best way to prevent RAM OOM errors is to split every big tile into smaller ones and save them on a disk. After that, make inference using the small crops and aggregate the result. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2237927": "I have taken middle 7 slices out of all the 65 slices and cut off the image 3d volume into 224x224 image slices but whenever I go for image 2 with seresnext model my RAM memory is full can someone provide any advice what shall I do",
    "2238371": "There are several tips I'm using in data loading for this competition:\n\n**- Preallocate memory instead of concatenation**\n`np.concatenate` or any other method to retrieve a single array from a list of many small arrays uses some overhead memory. The best thing to deal with it is to pre-allocate the array instead:\n```python\ndef load_tiles_efficient_way():\n    tile_idxs = get_tile_idxs()\n    data = np.array((len(tile_idxs), Config.TILE_SIZE, Config.TILE_SIZE, Config.N_SLICES), dtype=np.uint8)\n    for i, slice_path in enumerate(slice_paths):\n        slice_data = load_slice(slice_path)\n        data[..., i] = split_slice(slice_data, tile_idxs)\n    return data\n\ndef load_tiles_bad_way():\n    tile_idxs = get_tile_idxs()\n    tiles = []\n    for slice_path in slice_paths:\n        slice_data = load_slice(slice_path)\n        tiles.append(split_slice(slice_data, tile_idxs))\n    data = np.concatenate(tiles, axis=0))  # Uses more memory to hold both `tiles` list and `data` array\n    return data\n```\n\n**- Assign proper `dtype` for the volume `np.array` to reduce memory usage**\nThe difference between 16-bit values used in raw tiff files and 8-bit values shouldn't affect the model's performance too much.\nSo, you could store the volume data as `np.uint8` in the `[0-255]` range and convert it to `[0-1]` float in the `__getitem__` method of your `Dataset` class to reduce the memory usage constantly consumed by volume data.\n\n**- Don't store tiles outside the mask**\nYou could incorporate some logic to skip the tiles outside the `mask.png` if you didn't implement it yet. It will save you a serious amount of memory since a big part of the images is filled with zeros and doesn't represent anything useful in this competition.\n\n\n---\n\nOther advice could help someone:\n\n**- Use a smaller number of slices**\nFor those guys who decided to use all 65 slices, it would take you too much memory, and it's not possible to fit the whole dataset into your RAM, specifically on kaggle. On my setup, the whole dataset with all the tips above but with all 65 slices required at least 22 GB of RAM.\n\n**- Resize the images**\nThere are plenty of public notebooks that use that trick in different ways. But in general, you can resize the initial images to reduce memory usage.",
    "2239768": "Is that a problem with the RAM or vRAM (attached to GPU)? \n\n@r1chardson gives a lot of great advice already. I would add: try a smaller backbone and/or smaller batch size if it is a vRAM issue and try gradient accumulation to have a bigger \"virtual\" batch size.",
    "2240085": "In my experiments, the best way to prevent RAM OOM errors is to split every big tile into smaller ones and save them on a disk. After that, make inference using the small crops and aggregate the result."
  },
  "source": "meta"
}