{
  "id": 169344,
  "title": "DICOM Files Memory Leak",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/169344",
  "author_name": "Gunes Evitan",
  "post_date": "2020-07-23T15:32:20.362000",
  "votes": 15,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I noticed my memory consumption is gradually increasing while I iterate over the DICOM files. I tested the code below to see which part is causing the leak so I commented the parts that are accessing the fields of DICOM files. I think <code>dcm_read</code> is causing this leak.</p>\n\n<p><img src=\"https://i.ibb.co/RBwK1XJ/leak.jpg\" alt=\"\"></p>\n\n<p>Anyone experienced a similar memory issue?</p>",
  "messages": [
    {
      "id": 942112,
      "postDate": "2020-07-23T15:32:20.363Z",
      "content": "<p>I noticed my memory consumption is gradually increasing while I iterate over the DICOM files. I tested the code below to see which part is causing the leak so I commented the parts that are accessing the fields of DICOM files. I think <code>dcm_read</code> is causing this leak.</p>\n\n<p><img src=\"https://i.ibb.co/RBwK1XJ/leak.jpg\" alt=\"\"></p>\n\n<p>Anyone experienced a similar memory issue?</p>",
      "rawMarkdown": "I noticed my memory consumption is gradually increasing while I iterate over the DICOM files. I tested the code below to see which part is causing the leak so I commented the parts that are accessing the fields of DICOM files. I think `dcm_read` is causing this leak.\n\n![](https://i.ibb.co/RBwK1XJ/leak.jpg)\n\nAnyone experienced a similar memory issue?",
      "votes": 12
    },
    {
      "id": 966341,
      "postDate": "2020-08-11T10:45:11.977Z",
      "content": "<p>Just FYI, you can do <code>pydicom.dcmread(path, stop_before_pixels=True)</code> to read only the metadata into memory. Of course, this isn't useful if you actually want the pixel data. But if you're just converting the metadata into a CSV file it will save some memory.</p>",
      "rawMarkdown": "Just FYI, you can do `pydicom.dcmread(path, stop_before_pixels=True)` to read only the metadata into memory. Of course, this isn't useful if you actually want the pixel data. But if you're just converting the metadata into a CSV file it will save some memory.",
      "votes": 4,
      "replies": [
        {
          "id": 968214,
          "postDate": "2020-08-12T19:19:16.387Z",
          "content": "<p>Thanks, this is very useful.</p>",
          "rawMarkdown": "Thanks, this is very useful."
        }
      ]
    },
    {
      "id": 942213,
      "postDate": "2020-07-23T16:36:47.357Z",
      "content": "<p>Yes, I've had some trouble with memory usage as well, but it's not crashing my notebook anymore now that I'm mostly done editing the part of my code that handles image data, so I suspect there might be significant duplication happening under the hood if you're reading the same DCM files repeatedly (even if you use a <em>with</em> statement). When trying to diagnose the problem <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/167086\">I also found that the memory usage appears to be mysteriously inflated in certain circumstances</a>.</p>",
      "rawMarkdown": "Yes, I've had some trouble with memory usage as well, but it's not crashing my notebook anymore now that I'm mostly done editing the part of my code that handles image data, so I suspect there might be significant duplication happening under the hood if you're reading the same DCM files repeatedly (even if you use a *with* statement). When trying to diagnose the problem [I also found that the memory usage appears to be mysteriously inflated in certain circumstances](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/167086).",
      "votes": 1,
      "replies": [
        {
          "id": 942340,
          "postDate": "2020-07-23T17:48:22.713Z",
          "content": "<p>That is correct. Even reading the same file leaks memory. I tried using <code>with</code> statement, deleting the file and calling <code>gc.collect()</code>, but none of them worked.</p>\n\n<p>I found an <a href=\"https://github.com/pydicom/pydicom/issues/994\">issue</a> similar to this but it occurs when you access the fields.</p>",
          "rawMarkdown": "That is correct. Even reading the same file leaks memory. I tried using `with` statement, deleting the file and calling `gc.collect()`, but none of them worked.\n\nI found an [issue](https://github.com/pydicom/pydicom/issues/994) similar to this but it occurs when you access the fields."
        },
        {
          "id": 946787,
          "postDate": "2020-07-26T20:42:07.660Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 947285,
          "postDate": "2020-07-27T07:02:46.003Z",
          "content": "<p>Sadly, I haven't find any solution yet. I also tried the deprecated <code>pydicom.read_file</code> method and it has the same behavior.</p>",
          "rawMarkdown": "Sadly, I haven't find any solution yet. I also tried the deprecated `pydicom.read_file` method and it has the same behavior."
        },
        {
          "id": 947298,
          "postDate": "2020-07-27T07:12:45.690Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 942157,
      "postDate": "2020-07-23T16:02:16.593Z",
      "content": "<p>Yes I too had this issue.. my kernel crashed the first time while I iterated over the whole directory although I was using <code>pydicom.read_file</code> and it happened because I was storing each scan in a list so eventually the size of the list got too big and it ran out of memory..\nI corrected it by reducing the size of int in the pixel array.. I used <strong>int16</strong> which solved the problem\nhere is my kernel <a href=\"https://www.kaggle.com/zainahmad/preprocessing-the-dicom-data\">preprocessing dicom data</a></p>",
      "rawMarkdown": "Yes I too had this issue.. my kernel crashed the first time while I iterated over the whole directory although I was using `pydicom.read_file` and it happened because I was storing each scan in a list so eventually the size of the list got too big and it ran out of memory..\nI corrected it by reducing the size of int in the pixel array.. I used **int16** which solved the problem\nhere is my kernel [preprocessing dicom data](https://www.kaggle.com/zainahmad/preprocessing-the-dicom-data)",
      "votes": 1
    },
    {
      "id": 942548,
      "postDate": "2020-07-23T20:05:05.837Z",
      "content": "<p>Thnks <a href=\"/gunesevitan\">@gunesevitan</a>  to shed more light on it. I was stuck in my first CNN experiments till i change the IMAGE_SIZE. I understood now the source of the issue</p>",
      "rawMarkdown": "Thnks @gunesevitan  to shed more light on it. I was stuck in my first CNN experiments till i change the IMAGE_SIZE. I understood now the source of the issue",
      "votes": 2
    },
    {
      "id": 1057608,
      "postDate": "2020-10-22T19:47:33.763Z",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> Did you end up figuring this out? It's killing me in RSNA-PE :( I also found this <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106622\" target=\"_blank\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106622</a></p>",
      "rawMarkdown": "@gunesevitan Did you end up figuring this out? It's killing me in RSNA-PE :( I also found this https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106622",
      "replies": [
        {
          "id": 1057627,
          "postDate": "2020-10-22T20:00:32.563Z",
          "content": "<p>Sorry, I wasn't able to solve this problem. :/</p>",
          "rawMarkdown": "Sorry, I wasn't able to solve this problem. :/"
        },
        {
          "id": 1057638,
          "postDate": "2020-10-22T20:13:16.320Z",
          "content": "<p>That's alright: may be onto something here <a href=\"https://github.com/tqdm/tqdm/issues/746\" target=\"_blank\">https://github.com/tqdm/tqdm/issues/746</a></p>",
          "rawMarkdown": "That's alright: may be onto something here https://github.com/tqdm/tqdm/issues/746"
        }
      ]
    },
    {
      "id": 1028465,
      "postDate": "2020-09-26T21:50:37.333Z",
      "content": "<p>Hey did you get this issue resolved? I am getting similar using while reading the Dicom file using pydicom in my project.</p>",
      "rawMarkdown": "Hey did you get this issue resolved? I am getting similar using while reading the Dicom file using pydicom in my project."
    },
    {
      "id": 947625,
      "postDate": "2020-07-27T11:35:18.677Z",
      "content": "<p>One that's worked for me after suffering with the same problem is getting rid of tqdm. I need to refind the reference that I used but hopefully this helps some of you!</p>",
      "rawMarkdown": "One that's worked for me after suffering with the same problem is getting rid of tqdm. I need to refind the reference that I used but hopefully this helps some of you!"
    },
    {
      "id": 947543,
      "postDate": "2020-07-27T10:31:53.533Z",
      "content": "<p>I am not sure if this is the case, but most of the memory in DICOM is the actual image bytes, all metadata bytes (also stored in DICOM) combined takes less memory, so pydicom.dcmread actually make \"lazy\" computations, it only loads image bytes if you access them, e.g.</p>\n\n<p><code>img = pydicom.dcmread(\"img.dcm\")</code> &lt;- only consume small amount of memory\n<code>img.pixel_array</code> &lt;- will consume much more</p>\n\n<p>if your only intention is to read metainformation from DICOM header you could do it once and store it in a separate csv file, see for example this notebook <a href=\"https://www.kaggle.com/akurmukov/collect-dicoms-metadata-interactive-visualization\">https://www.kaggle.com/akurmukov/collect-dicoms-metadata-interactive-visualization</a></p>",
      "rawMarkdown": "I am not sure if this is the case, but most of the memory in DICOM is the actual image bytes, all metadata bytes (also stored in DICOM) combined takes less memory, so pydicom.dcmread actually make \"lazy\" computations, it only loads image bytes if you access them, e.g.\n\n`img = pydicom.dcmread(\"img.dcm\")` &lt;- only consume small amount of memory\n`img.pixel_array` &lt;- will consume much more\n\nif your only intention is to read metainformation from DICOM header you could do it once and store it in a separate csv file, see for example this notebook https://www.kaggle.com/akurmukov/collect-dicoms-metadata-interactive-visualization",
      "replies": [
        {
          "id": 947560,
          "postDate": "2020-07-27T10:46:53.687Z",
          "content": "<p>The code I shared is only accessing the metadata, not pixel array.</p>",
          "rawMarkdown": "The code I shared is only accessing the metadata, not pixel array."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 966341,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2020-08-11T10:45:11.977000",
      "content": "<p>Just FYI, you can do <code>pydicom.dcmread(path, stop_before_pixels=True)</code> to read only the metadata into memory. Of course, this isn't useful if you actually want the pixel data. But if you're just converting the metadata into a CSV file it will save some memory.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 968214,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2020-08-12T19:19:16.387000",
          "content": "<p>Thanks, this is very useful.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 942213,
      "author_name": "Mathieu Beaudoin",
      "author_url": "",
      "post_date": "2020-07-23T16:36:47.357000",
      "content": "<p>Yes, I've had some trouble with memory usage as well, but it's not crashing my notebook anymore now that I'm mostly done editing the part of my code that handles image data, so I suspect there might be significant duplication happening under the hood if you're reading the same DCM files repeatedly (even if you use a <em>with</em> statement). When trying to diagnose the problem <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/167086\">I also found that the memory usage appears to be mysteriously inflated in certain circumstances</a>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 942340,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2020-07-23T17:48:22.713000",
          "content": "<p>That is correct. Even reading the same file leaks memory. I tried using <code>with</code> statement, deleting the file and calling <code>gc.collect()</code>, but none of them worked.</p>\n\n<p>I found an <a href=\"https://github.com/pydicom/pydicom/issues/994\">issue</a> similar to this but it occurs when you access the fields.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946787,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-26T20:42:07.660000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 947285,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2020-07-27T07:02:46.003000",
          "content": "<p>Sadly, I haven't find any solution yet. I also tried the deprecated <code>pydicom.read_file</code> method and it has the same behavior.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 947298,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-27T07:12:45.690000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 942157,
      "author_name": "Zain Ahmad",
      "author_url": "",
      "post_date": "2020-07-23T16:02:16.593000",
      "content": "<p>Yes I too had this issue.. my kernel crashed the first time while I iterated over the whole directory although I was using <code>pydicom.read_file</code> and it happened because I was storing each scan in a list so eventually the size of the list got too big and it ran out of memory..\nI corrected it by reducing the size of int in the pixel array.. I used <strong>int16</strong> which solved the problem\nhere is my kernel <a href=\"https://www.kaggle.com/zainahmad/preprocessing-the-dicom-data\">preprocessing dicom data</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 942548,
      "author_name": "Ulrich G.",
      "author_url": "",
      "post_date": "2020-07-23T20:05:05.837000",
      "content": "<p>Thnks <a href=\"/gunesevitan\">@gunesevitan</a>  to shed more light on it. I was stuck in my first CNN experiments till i change the IMAGE_SIZE. I understood now the source of the issue</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1057608,
      "author_name": "Alexander Soare",
      "author_url": "",
      "post_date": "2020-10-22T19:47:33.763000",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> Did you end up figuring this out? It's killing me in RSNA-PE :( I also found this <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106622\" target=\"_blank\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106622</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1057627,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2020-10-22T20:00:32.563000",
          "content": "<p>Sorry, I wasn't able to solve this problem. :/</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057638,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2020-10-22T20:13:16.320000",
          "content": "<p>That's alright: may be onto something here <a href=\"https://github.com/tqdm/tqdm/issues/746\" target=\"_blank\">https://github.com/tqdm/tqdm/issues/746</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1028465,
      "author_name": "SumanSudhir",
      "author_url": "",
      "post_date": "2020-09-26T21:50:37.333000",
      "content": "<p>Hey did you get this issue resolved? I am getting similar using while reading the Dicom file using pydicom in my project.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 947625,
      "author_name": "jameschapman19",
      "author_url": "",
      "post_date": "2020-07-27T11:35:18.677000",
      "content": "<p>One that's worked for me after suffering with the same problem is getting rid of tqdm. I need to refind the reference that I used but hopefully this helps some of you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 947543,
      "author_name": "akurmukov",
      "author_url": "",
      "post_date": "2020-07-27T10:31:53.533000",
      "content": "<p>I am not sure if this is the case, but most of the memory in DICOM is the actual image bytes, all metadata bytes (also stored in DICOM) combined takes less memory, so pydicom.dcmread actually make \"lazy\" computations, it only loads image bytes if you access them, e.g.</p>\n\n<p><code>img = pydicom.dcmread(\"img.dcm\")</code> &lt;- only consume small amount of memory\n<code>img.pixel_array</code> &lt;- will consume much more</p>\n\n<p>if your only intention is to read metainformation from DICOM header you could do it once and store it in a separate csv file, see for example this notebook <a href=\"https://www.kaggle.com/akurmukov/collect-dicoms-metadata-interactive-visualization\">https://www.kaggle.com/akurmukov/collect-dicoms-metadata-interactive-visualization</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 947560,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2020-07-27T10:46:53.687000",
          "content": "<p>The code I shared is only accessing the metadata, not pixel array.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "942112": "I noticed my memory consumption is gradually increasing while I iterate over the DICOM files. I tested the code below to see which part is causing the leak so I commented the parts that are accessing the fields of DICOM files. I think `dcm_read` is causing this leak.\n\n![](https://i.ibb.co/RBwK1XJ/leak.jpg)\n\nAnyone experienced a similar memory issue?",
    "966341": "Just FYI, you can do `pydicom.dcmread(path, stop_before_pixels=True)` to read only the metadata into memory. Of course, this isn't useful if you actually want the pixel data. But if you're just converting the metadata into a CSV file it will save some memory.",
    "942213": "Yes, I've had some trouble with memory usage as well, but it's not crashing my notebook anymore now that I'm mostly done editing the part of my code that handles image data, so I suspect there might be significant duplication happening under the hood if you're reading the same DCM files repeatedly (even if you use a *with* statement). When trying to diagnose the problem [I also found that the memory usage appears to be mysteriously inflated in certain circumstances](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/167086).",
    "942157": "Yes I too had this issue.. my kernel crashed the first time while I iterated over the whole directory although I was using `pydicom.read_file` and it happened because I was storing each scan in a list so eventually the size of the list got too big and it ran out of memory..\nI corrected it by reducing the size of int in the pixel array.. I used **int16** which solved the problem\nhere is my kernel [preprocessing dicom data](https://www.kaggle.com/zainahmad/preprocessing-the-dicom-data)",
    "942548": "Thnks @gunesevitan  to shed more light on it. I was stuck in my first CNN experiments till i change the IMAGE_SIZE. I understood now the source of the issue",
    "1057608": "@gunesevitan Did you end up figuring this out? It's killing me in RSNA-PE :( I also found this https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106622",
    "1028465": "Hey did you get this issue resolved? I am getting similar using while reading the Dicom file using pydicom in my project.",
    "947625": "One that's worked for me after suffering with the same problem is getting rid of tqdm. I need to refind the reference that I used but hopefully this helps some of you!",
    "947543": "I am not sure if this is the case, but most of the memory in DICOM is the actual image bytes, all metadata bytes (also stored in DICOM) combined takes less memory, so pydicom.dcmread actually make \"lazy\" computations, it only loads image bytes if you access them, e.g.\n\n`img = pydicom.dcmread(\"img.dcm\")` &lt;- only consume small amount of memory\n`img.pixel_array` &lt;- will consume much more\n\nif your only intention is to read metainformation from DICOM header you could do it once and store it in a separate csv file, see for example this notebook https://www.kaggle.com/akurmukov/collect-dicoms-metadata-interactive-visualization"
  }
}