{
  "id": 273096,
  "title": "Getting issues with memory while converting all images to numpy array and not able to execute in kaggle notebook instance. is there any alternative to get rid of memory issues",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273096",
  "author_name": "",
  "post_date": "2021-09-19T06:21:36.373496900Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Getting issues with memory while converting all images to numpy array and not able to execute in kaggle notebook instance. is there any alternative to get rid of memory issues</p>",
  "messages": [
    {
      "id": "1517046",
      "postDate": "09/19/2021 06:21:36",
      "content": "<p>Getting issues with memory while converting all images to numpy array and not able to execute in kaggle notebook instance. is there any alternative to get rid of memory issues</p>",
      "rawMarkdown": "Getting issues with memory while converting all images to numpy array and not able to execute in kaggle notebook instance. is there any alternative to get rid of memory issues",
      "votes": null
    },
    {
      "id": "1517117",
      "postDate": "09/19/2021 08:06:01",
      "content": "<ol>\n<li>can use .npz instead of npy. compress them np.savez_compressed.</li>\n<li>load a numpy dataset that someone else created before you.</li>\n</ol>",
      "rawMarkdown": "1. can use .npz instead of npy. compress them np.savez_compressed.\n2. load a numpy dataset that someone else created before you.",
      "votes": null
    },
    {
      "id": "1517337",
      "postDate": "09/19/2021 13:33:28",
      "content": "<ol>\n<li>Specify the numpy array as np.float32 works for data with size 128x128x128. I have tried 256x256x256 but also have memory issue </li>\n<li>Delete variables that you don't need anymore.<br>\nEx: the train data csv once you have extracted the IDs and labels you need. <br>\nimport gc<br>\ndel train_data<br>\ngc.collect()</li>\n</ol>",
      "rawMarkdown": "1. Specify the numpy array as np.float32 works for data with size 128x128x128. I have tried 256x256x256 but also have memory issue \n2. Delete variables that you don't need anymore.\nEx: the train data csv once you have extracted the IDs and labels you need. \nimport gc\ndel train_data\ngc.collect()",
      "votes": null
    },
    {
      "id": "1518310",
      "postDate": "09/20/2021 14:53:33",
      "content": "<p>Use np.array.nbytes to know the no of bytes used by your NumPy arrays. If they are over the limit then try to use a smaller image size, a smaller number of images from each folder, and don't forget you need memory also to train your model, use smaller batch size and bigger stride for CNNs. And don't try to fit all of your images into memory at once before training or testing, use data generators.</p>",
      "rawMarkdown": "Use np.array.nbytes to know the no of bytes used by your NumPy arrays. If they are over the limit then try to use a smaller image size, a smaller number of images from each folder, and don't forget you need memory also to train your model, use smaller batch size and bigger stride for CNNs. And don't try to fit all of your images into memory at once before training or testing, use data generators.",
      "votes": null
    },
    {
      "id": "1518518",
      "postDate": "09/20/2021 17:52:30",
      "content": "<p>Aside from the methods mentioned in this thread, I highly recommend using mixed precision. It can be easily turned on in both tensorflow and pytorch. It saves a lot of memory! </p>",
      "rawMarkdown": "Aside from the methods mentioned in this thread, I highly recommend using mixed precision. It can be easily turned on in both tensorflow and pytorch. It saves a lot of memory!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1517117,
      "author_name": "zaakciiru",
      "author_url": "",
      "post_date": "09/19/2021 08:06:01",
      "content": "<ol>\n<li>can use .npz instead of npy. compress them np.savez_compressed.</li>\n<li>load a numpy dataset that someone else created before you.</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1517337,
      "author_name": "stelis",
      "author_url": "",
      "post_date": "09/19/2021 13:33:28",
      "content": "<ol>\n<li>Specify the numpy array as np.float32 works for data with size 128x128x128. I have tried 256x256x256 but also have memory issue </li>\n<li>Delete variables that you don't need anymore.<br>\nEx: the train data csv once you have extracted the IDs and labels you need. <br>\nimport gc<br>\ndel train_data<br>\ngc.collect()</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1518310,
      "author_name": "susnato",
      "author_url": "",
      "post_date": "09/20/2021 14:53:33",
      "content": "<p>Use np.array.nbytes to know the no of bytes used by your NumPy arrays. If they are over the limit then try to use a smaller image size, a smaller number of images from each folder, and don't forget you need memory also to train your model, use smaller batch size and bigger stride for CNNs. And don't try to fit all of your images into memory at once before training or testing, use data generators.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1518518,
      "author_name": "mikecho",
      "author_url": "",
      "post_date": "09/20/2021 17:52:30",
      "content": "<p>Aside from the methods mentioned in this thread, I highly recommend using mixed precision. It can be easily turned on in both tensorflow and pytorch. It saves a lot of memory! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1517046": "Getting issues with memory while converting all images to numpy array and not able to execute in kaggle notebook instance. is there any alternative to get rid of memory issues",
    "1517117": "1. can use .npz instead of npy. compress them np.savez_compressed.\n2. load a numpy dataset that someone else created before you.",
    "1517337": "1. Specify the numpy array as np.float32 works for data with size 128x128x128. I have tried 256x256x256 but also have memory issue \n2. Delete variables that you don't need anymore.\nEx: the train data csv once you have extracted the IDs and labels you need. \nimport gc\ndel train_data\ngc.collect()",
    "1518310": "Use np.array.nbytes to know the no of bytes used by your NumPy arrays. If they are over the limit then try to use a smaller image size, a smaller number of images from each folder, and don't forget you need memory also to train your model, use smaller batch size and bigger stride for CNNs. And don't try to fit all of your images into memory at once before training or testing, use data generators.",
    "1518518": "Aside from the methods mentioned in this thread, I highly recommend using mixed precision. It can be easily turned on in both tensorflow and pytorch. It saves a lot of memory!"
  },
  "source": "meta"
}