{
  "id": 126623,
  "title": "Memory leaks in Kaggle's kernel",
  "url": "/competitions/deepfake-detection-challenge/discussion/126623",
  "author_name": "",
  "post_date": "2020-01-18T22:05:26.629019300Z",
  "votes": 6,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hi guys</p>\n\n<p>Is anyone observing memory leaks in Kaggle's kernel? For me it always happens with OpenCV (VideoCapture) no matter what. I've already tried using function from public kernel and it didn't work either. Funny thing is the same code run fine on my local machine.</p>",
  "messages": [
    {
      "id": "722661",
      "postDate": "01/18/2020 22:05:26",
      "content": "<p>Hi guys</p>\n\n<p>Is anyone observing memory leaks in Kaggle's kernel? For me it always happens with OpenCV (VideoCapture) no matter what. I've already tried using function from public kernel and it didn't work either. Funny thing is the same code run fine on my local machine.</p>",
      "rawMarkdown": "Hi guys\n\nIs anyone observing memory leaks in Kaggle's kernel? For me it always happens with OpenCV (VideoCapture) no matter what. I've already tried using function from public kernel and it didn't work either. Funny thing is the same code run fine on my local machine.",
      "votes": null
    },
    {
      "id": "722683",
      "postDate": "01/18/2020 23:14:51",
      "content": "<p>Yes, I had similar issues:\n<a href=\"https://www.kaggle.com/timesler/guide-to-mtcnn-in-facenet-pytorch#710202\">https://www.kaggle.com/timesler/guide-to-mtcnn-in-facenet-pytorch#710202</a></p>\n\n<p>And one time (or maybe two), I got out of memory with GPU on Kaggle for some reasons. Just running the same code again worked. Weird. </p>",
      "rawMarkdown": "Yes, I had similar issues:\nhttps://www.kaggle.com/timesler/guide-to-mtcnn-in-facenet-pytorch#710202\n\nAnd one time (or maybe two), I got out of memory with GPU on Kaggle for some reasons. Just running the same code again worked. Weird.",
      "votes": null
    },
    {
      "id": "722996",
      "postDate": "01/19/2020 11:28:27",
      "content": "<p>Thanks. Just out of curiosity, what is the intuition of using copy to fix memory leaks?</p>",
      "rawMarkdown": "Thanks. Just out of curiosity, what is the intuition of using copy to fix memory leaks?",
      "votes": null
    },
    {
      "id": "723018",
      "postDate": "01/19/2020 11:55:36",
      "content": "<p>I've enabled python tracemalloc to track the leak and memory increase was coming from each image (either CV2 or PIL) on the <code>np.append(small_face)</code>. The full image was retained in memory instead of the small face. So I tried many things to free it and <code>.copy()</code> worked. It only happened with GPU, no leak with CPU. It's like image was retained by a GPU side effect. The root cause might be obvious for Pytorch/TF/GPU experts but I'm not expert.</p>",
      "rawMarkdown": "I've enabled python tracemalloc to track the leak and memory increase was coming from each image (either CV2 or PIL) on the `np.append(small_face)`. The full image was retained in memory instead of the small face. So I tried many things to free it and `.copy()` worked. It only happened with GPU, no leak with CPU. It's like image was retained by a GPU side effect. The root cause might be obvious for Pytorch/TF/GPU experts but I'm not expert.",
      "votes": null
    },
    {
      "id": "723095",
      "postDate": "01/19/2020 14:04:34",
      "content": "<p>I see, thanks. But for me it didn't work. I'm no expert in profilling memory, but I've tried tracemalloc and pympler and neither worked, this is really frustrating :/</p>",
      "rawMarkdown": "I see, thanks. But for me it didn't work. I'm no expert in profilling memory, but I've tried tracemalloc and pympler and neither worked, this is really frustrating :/",
      "votes": null
    },
    {
      "id": "723196",
      "postDate": "01/19/2020 16:17:02",
      "content": "<p>Not sure what you were doing, but if you do <code>small_face = big_face[t:b, l:r, :]</code> then you're making a slice into the original array and the system will need to keep the entire original array around. By doing <code>.copy()</code>, it only copies the pixels from the slice into a new array, and the original can be freed.</p>",
      "rawMarkdown": "Not sure what you were doing, but if you do `small_face = big_face[t:b, l:r, :]` then you're making a slice into the original array and the system will need to keep the entire original array around. By doing `.copy()`, it only copies the pixels from the slice into a new array, and the original can be freed.",
      "votes": null
    },
    {
      "id": "725869",
      "postDate": "01/22/2020 15:00:50",
      "content": "<p>I seem to be having the same issue. I'm only using (VideoCapture) to extract one frame at a time to then extract features from it, but memory use keeps increasing drmatically even though I'm not saving the picture to any array.</p>",
      "rawMarkdown": "I seem to be having the same issue. I'm only using (VideoCapture) to extract one frame at a time to then extract features from it, but memory use keeps increasing drmatically even though I'm not saving the picture to any array.",
      "votes": null
    },
    {
      "id": "725889",
      "postDate": "01/22/2020 15:23:51",
      "content": "<p>Try submitting the kernel anyway. For me, the submission worked although the UI was clearly showing a memory leak during kernel's interactive mode </p>",
      "rawMarkdown": "Try submitting the kernel anyway. For me, the submission worked although the UI was clearly showing a memory leak during kernel's interactive mode",
      "votes": null
    },
    {
      "id": "725930",
      "postDate": "01/22/2020 16:18:17",
      "content": "<p>I just noticed the same thing happening in my kernel. Looping through all the videos reading a single frame per video will bump memory usage to 1.3 GB and it never gets reclaimed. </p>\n\n<p>In my inference kernel memory goes up to 15.9 GB -- eek! It does work for making a submission, but it looks like I've got some memory debugging to do. Some of this seems to be caused by OpenCV's VideoCapture, but it doesn't explain everything...</p>",
      "rawMarkdown": "I just noticed the same thing happening in my kernel. Looping through all the videos reading a single frame per video will bump memory usage to 1.3 GB and it never gets reclaimed. \n\nIn my inference kernel memory goes up to 15.9 GB -- eek! It does work for making a submission, but it looks like I've got some memory debugging to do. Some of this seems to be caused by OpenCV's VideoCapture, but it doesn't explain everything...",
      "votes": null
    },
    {
      "id": "726577",
      "postDate": "01/23/2020 05:12:21",
      "content": "<p>Similar Issue:\nJust added x=x/x.max() and ran out of memory. Even though it is not even close to 12 GB.</p>",
      "rawMarkdown": "Similar Issue:\nJust added x=x/x.max() and ran out of memory. Even though it is not even close to 12 GB.",
      "votes": null
    },
    {
      "id": "726932",
      "postDate": "01/23/2020 10:18:41",
      "content": "<p>For me, debugging memory was in vain. I just trusted that the problem were in the UI and it wouldn't occur in the private dataset. At least so far, I'm not having any problems in my submissions so it is probably an UI bug</p>",
      "rawMarkdown": "For me, debugging memory was in vain. I just trusted that the problem were in the UI and it wouldn't occur in the private dataset. At least so far, I'm not having any problems in my submissions so it is probably an UI bug",
      "votes": null
    },
    {
      "id": "726953",
      "postDate": "01/23/2020 10:35:26",
      "content": "<p>I'm thinking my issue is a PyTorch memory bug (?) with running on the CPU. The memory usage is stable when using the GPU.</p>",
      "rawMarkdown": "I'm thinking my issue is a PyTorch memory bug (?) with running on the CPU. The memory usage is stable when using the GPU.",
      "votes": null
    },
    {
      "id": "726965",
      "postDate": "01/23/2020 10:44:14",
      "content": "<p>Hmm, there is also a real memory real in pytorch if you use multiprocessing and have lists, tuples or ndarray (of object type) in the Dataset. Take a look <a href=\"https://github.com/pytorch/pytorch/issues/13246\">https://github.com/pytorch/pytorch/issues/13246</a></p>",
      "rawMarkdown": "Hmm, there is also a real memory real in pytorch if you use multiprocessing and have lists, tuples or ndarray (of object type) in the Dataset. Take a look https://github.com/pytorch/pytorch/issues/13246",
      "votes": null
    },
    {
      "id": "728529",
      "postDate": "01/24/2020 21:22:38",
      "content": "<p>Do you always call the <code>.release()</code> function on VideoReader for each video? ...maybe this might be the reason of your leak</p>",
      "rawMarkdown": "Do you always call the `.release()` function on VideoReader for each video? ...maybe this might be the reason of your leak",
      "votes": null
    },
    {
      "id": "728562",
      "postDate": "01/24/2020 22:42:47",
      "content": "<p>I do call <code>.release()</code>. 😄 </p>\n\n<p>A simple test that loads all test videos in a loop eats up about 1.5 GB of memory on Kaggle (that never gets returned) while on my local machine it's only about 100 MB. I've heard it mentioned in the forums here that OpenCV on Kaggle has a memory leak issue and it seems plausible.</p>\n\n<p>As for my more serious memory bug that eats up all RAM when using PyTorch, this appears to happen if you change the batch size when doing inference on the CPU. It took me 3 days to figure out that the varying batch size was the culprit here. (<a href=\"https://github.com/pytorch/pytorch/issues/32596\">I made a bug report in case you're curious.</a>)</p>",
      "rawMarkdown": "I do call `.release()`. 😄 \n\nA simple test that loads all test videos in a loop eats up about 1.5 GB of memory on Kaggle (that never gets returned) while on my local machine it's only about 100 MB. I've heard it mentioned in the forums here that OpenCV on Kaggle has a memory leak issue and it seems plausible.\n\nAs for my more serious memory bug that eats up all RAM when using PyTorch, this appears to happen if you change the batch size when doing inference on the CPU. It took me 3 days to figure out that the varying batch size was the culprit here. ([I made a bug report in case you're curious.](https://github.com/pytorch/pytorch/issues/32596))",
      "votes": null
    },
    {
      "id": "728681",
      "postDate": "01/25/2020 04:48:49",
      "content": "<p>Seems to be a problem with opencv itself: <a href=\"https://github.com/opencv/opencv/issues/5715#issuecomment-500116136\">https://github.com/opencv/opencv/issues/5715#issuecomment-500116136</a>\nHaven't figured out a way to fix it yet in python.</p>",
      "rawMarkdown": "Seems to be a problem with opencv itself: https://github.com/opencv/opencv/issues/5715#issuecomment-500116136\nHaven't figured out a way to fix it yet in python.",
      "votes": null
    },
    {
      "id": "729185",
      "postDate": "01/25/2020 21:44:22",
      "content": "<p>Yes, I tried deleting all of the local variables and calling garbage collection within the loops (including any release functionality), but upon returning from the function, the memory was still allocated. Capturing 11 frames apiece for 400 videos drove the memory up to ~16 GB.</p>",
      "rawMarkdown": "Yes, I tried deleting all of the local variables and calling garbage collection within the loops (including any release functionality), but upon returning from the function, the memory was still allocated. Capturing 11 frames apiece for 400 videos drove the memory up to ~16 GB.",
      "votes": null
    },
    {
      "id": "729535",
      "postDate": "01/26/2020 10:27:50",
      "content": "<blockquote>\n  <p>Capturing 11 frames apiece for 400 videos drove the memory up to ~16 GB.</p>\n</blockquote>\n\n<p>That's much more severe than the memory leak I found with OpenCV... It sounds like something else is going wrong in your loop as well.</p>",
      "rawMarkdown": "&gt; Capturing 11 frames apiece for 400 videos drove the memory up to ~16 GB.\n\nThat's much more severe than the memory leak I found with OpenCV... It sounds like something else is going wrong in your loop as well.",
      "votes": null
    },
    {
      "id": "778337",
      "postDate": "03/18/2020 11:01:22",
      "content": "<p>I have the same issue, how did any of you resolve this? if 400 videos drives memory to ~16GB how can you have enough memory for 4000 videos even if submission works or for training? </p>",
      "rawMarkdown": "I have the same issue, how did any of you resolve this? if 400 videos drives memory to ~16GB how can you have enough memory for 4000 videos even if submission works or for training?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 722683,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "01/18/2020 23:14:51",
      "content": "<p>Yes, I had similar issues:\n<a href=\"https://www.kaggle.com/timesler/guide-to-mtcnn-in-facenet-pytorch#710202\">https://www.kaggle.com/timesler/guide-to-mtcnn-in-facenet-pytorch#710202</a></p>\n\n<p>And one time (or maybe two), I got out of memory with GPU on Kaggle for some reasons. Just running the same code again worked. Weird. </p>",
      "votes": null,
      "replies": [
        {
          "id": 722996,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "01/19/2020 11:28:27",
          "content": "<p>Thanks. Just out of curiosity, what is the intuition of using copy to fix memory leaks?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 723018,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "01/19/2020 11:55:36",
          "content": "<p>I've enabled python tracemalloc to track the leak and memory increase was coming from each image (either CV2 or PIL) on the <code>np.append(small_face)</code>. The full image was retained in memory instead of the small face. So I tried many things to free it and <code>.copy()</code> worked. It only happened with GPU, no leak with CPU. It's like image was retained by a GPU side effect. The root cause might be obvious for Pytorch/TF/GPU experts but I'm not expert.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 723095,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "01/19/2020 14:04:34",
          "content": "<p>I see, thanks. But for me it didn't work. I'm no expert in profilling memory, but I've tried tracemalloc and pympler and neither worked, this is really frustrating :/</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 723196,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/19/2020 16:17:02",
          "content": "<p>Not sure what you were doing, but if you do <code>small_face = big_face[t:b, l:r, :]</code> then you're making a slice into the original array and the system will need to keep the entire original array around. By doing <code>.copy()</code>, it only copies the pixels from the slice into a new array, and the original can be freed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 725869,
      "author_name": "camyok",
      "author_url": "",
      "post_date": "01/22/2020 15:00:50",
      "content": "<p>I seem to be having the same issue. I'm only using (VideoCapture) to extract one frame at a time to then extract features from it, but memory use keeps increasing drmatically even though I'm not saving the picture to any array.</p>",
      "votes": null,
      "replies": [
        {
          "id": 725889,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "01/22/2020 15:23:51",
          "content": "<p>Try submitting the kernel anyway. For me, the submission worked although the UI was clearly showing a memory leak during kernel's interactive mode </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 725930,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/22/2020 16:18:17",
          "content": "<p>I just noticed the same thing happening in my kernel. Looping through all the videos reading a single frame per video will bump memory usage to 1.3 GB and it never gets reclaimed. </p>\n\n<p>In my inference kernel memory goes up to 15.9 GB -- eek! It does work for making a submission, but it looks like I've got some memory debugging to do. Some of this seems to be caused by OpenCV's VideoCapture, but it doesn't explain everything...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 726932,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "01/23/2020 10:18:41",
          "content": "<p>For me, debugging memory was in vain. I just trusted that the problem were in the UI and it wouldn't occur in the private dataset. At least so far, I'm not having any problems in my submissions so it is probably an UI bug</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 726953,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/23/2020 10:35:26",
          "content": "<p>I'm thinking my issue is a PyTorch memory bug (?) with running on the CPU. The memory usage is stable when using the GPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 726965,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "01/23/2020 10:44:14",
          "content": "<p>Hmm, there is also a real memory real in pytorch if you use multiprocessing and have lists, tuples or ndarray (of object type) in the Dataset. Take a look <a href=\"https://github.com/pytorch/pytorch/issues/13246\">https://github.com/pytorch/pytorch/issues/13246</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 726577,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "01/23/2020 05:12:21",
      "content": "<p>Similar Issue:\nJust added x=x/x.max() and ran out of memory. Even though it is not even close to 12 GB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 728529,
      "author_name": "dagnelies",
      "author_url": "",
      "post_date": "01/24/2020 21:22:38",
      "content": "<p>Do you always call the <code>.release()</code> function on VideoReader for each video? ...maybe this might be the reason of your leak</p>",
      "votes": null,
      "replies": [
        {
          "id": 728562,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/24/2020 22:42:47",
          "content": "<p>I do call <code>.release()</code>. 😄 </p>\n\n<p>A simple test that loads all test videos in a loop eats up about 1.5 GB of memory on Kaggle (that never gets returned) while on my local machine it's only about 100 MB. I've heard it mentioned in the forums here that OpenCV on Kaggle has a memory leak issue and it seems plausible.</p>\n\n<p>As for my more serious memory bug that eats up all RAM when using PyTorch, this appears to happen if you change the batch size when doing inference on the CPU. It took me 3 days to figure out that the varying batch size was the culprit here. (<a href=\"https://github.com/pytorch/pytorch/issues/32596\">I made a bug report in case you're curious.</a>)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 728681,
      "author_name": "teknas",
      "author_url": "",
      "post_date": "01/25/2020 04:48:49",
      "content": "<p>Seems to be a problem with opencv itself: <a href=\"https://github.com/opencv/opencv/issues/5715#issuecomment-500116136\">https://github.com/opencv/opencv/issues/5715#issuecomment-500116136</a>\nHaven't figured out a way to fix it yet in python.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 729185,
      "author_name": "mconway",
      "author_url": "",
      "post_date": "01/25/2020 21:44:22",
      "content": "<p>Yes, I tried deleting all of the local variables and calling garbage collection within the loops (including any release functionality), but upon returning from the function, the memory was still allocated. Capturing 11 frames apiece for 400 videos drove the memory up to ~16 GB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 729535,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/26/2020 10:27:50",
          "content": "<blockquote>\n  <p>Capturing 11 frames apiece for 400 videos drove the memory up to ~16 GB.</p>\n</blockquote>\n\n<p>That's much more severe than the memory leak I found with OpenCV... It sounds like something else is going wrong in your loop as well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 778337,
      "author_name": "ozzydeeeplearner",
      "author_url": "",
      "post_date": "03/18/2020 11:01:22",
      "content": "<p>I have the same issue, how did any of you resolve this? if 400 videos drives memory to ~16GB how can you have enough memory for 4000 videos even if submission works or for training? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "722661": "Hi guys\n\nIs anyone observing memory leaks in Kaggle's kernel? For me it always happens with OpenCV (VideoCapture) no matter what. I've already tried using function from public kernel and it didn't work either. Funny thing is the same code run fine on my local machine.",
    "722683": "Yes, I had similar issues:\nhttps://www.kaggle.com/timesler/guide-to-mtcnn-in-facenet-pytorch#710202\n\nAnd one time (or maybe two), I got out of memory with GPU on Kaggle for some reasons. Just running the same code again worked. Weird.",
    "722996": "Thanks. Just out of curiosity, what is the intuition of using copy to fix memory leaks?",
    "723018": "I've enabled python tracemalloc to track the leak and memory increase was coming from each image (either CV2 or PIL) on the `np.append(small_face)`. The full image was retained in memory instead of the small face. So I tried many things to free it and `.copy()` worked. It only happened with GPU, no leak with CPU. It's like image was retained by a GPU side effect. The root cause might be obvious for Pytorch/TF/GPU experts but I'm not expert.",
    "723095": "I see, thanks. But for me it didn't work. I'm no expert in profilling memory, but I've tried tracemalloc and pympler and neither worked, this is really frustrating :/",
    "723196": "Not sure what you were doing, but if you do `small_face = big_face[t:b, l:r, :]` then you're making a slice into the original array and the system will need to keep the entire original array around. By doing `.copy()`, it only copies the pixels from the slice into a new array, and the original can be freed.",
    "725869": "I seem to be having the same issue. I'm only using (VideoCapture) to extract one frame at a time to then extract features from it, but memory use keeps increasing drmatically even though I'm not saving the picture to any array.",
    "725889": "Try submitting the kernel anyway. For me, the submission worked although the UI was clearly showing a memory leak during kernel's interactive mode",
    "725930": "I just noticed the same thing happening in my kernel. Looping through all the videos reading a single frame per video will bump memory usage to 1.3 GB and it never gets reclaimed. \n\nIn my inference kernel memory goes up to 15.9 GB -- eek! It does work for making a submission, but it looks like I've got some memory debugging to do. Some of this seems to be caused by OpenCV's VideoCapture, but it doesn't explain everything...",
    "726577": "Similar Issue:\nJust added x=x/x.max() and ran out of memory. Even though it is not even close to 12 GB.",
    "726932": "For me, debugging memory was in vain. I just trusted that the problem were in the UI and it wouldn't occur in the private dataset. At least so far, I'm not having any problems in my submissions so it is probably an UI bug",
    "726953": "I'm thinking my issue is a PyTorch memory bug (?) with running on the CPU. The memory usage is stable when using the GPU.",
    "726965": "Hmm, there is also a real memory real in pytorch if you use multiprocessing and have lists, tuples or ndarray (of object type) in the Dataset. Take a look https://github.com/pytorch/pytorch/issues/13246",
    "728529": "Do you always call the `.release()` function on VideoReader for each video? ...maybe this might be the reason of your leak",
    "728562": "I do call `.release()`. 😄 \n\nA simple test that loads all test videos in a loop eats up about 1.5 GB of memory on Kaggle (that never gets returned) while on my local machine it's only about 100 MB. I've heard it mentioned in the forums here that OpenCV on Kaggle has a memory leak issue and it seems plausible.\n\nAs for my more serious memory bug that eats up all RAM when using PyTorch, this appears to happen if you change the batch size when doing inference on the CPU. It took me 3 days to figure out that the varying batch size was the culprit here. ([I made a bug report in case you're curious.](https://github.com/pytorch/pytorch/issues/32596))",
    "728681": "Seems to be a problem with opencv itself: https://github.com/opencv/opencv/issues/5715#issuecomment-500116136\nHaven't figured out a way to fix it yet in python.",
    "729185": "Yes, I tried deleting all of the local variables and calling garbage collection within the loops (including any release functionality), but upon returning from the function, the memory was still allocated. Capturing 11 frames apiece for 400 videos drove the memory up to ~16 GB.",
    "729535": "&gt; Capturing 11 frames apiece for 400 videos drove the memory up to ~16 GB.\n\nThat's much more severe than the memory leak I found with OpenCV... It sounds like something else is going wrong in your loop as well.",
    "778337": "I have the same issue, how did any of you resolve this? if 400 videos drives memory to ~16GB how can you have enough memory for 4000 videos even if submission works or for training?"
  },
  "source": "meta"
}