{
  "id": 124453,
  "title": "Notebook Exceeded Allowed Compute",
  "url": "/competitions/deepfake-detection-challenge/discussion/124453",
  "author_name": "",
  "post_date": "2020-01-04T07:59:42.780326900Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I trained a model based on Xception and MTCNN. My Notebook runs successfully on my local machine, and I guarantee that the model will run in less than nine hours. But when it comes to submission, this problem keeps coming up. Does anybody know the reason for this?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2Fe2e21cd7c43deb1b90fe9f7be54fec1b%2F1578124856714.jpg?generation=1578124879206163&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "710006",
      "postDate": "01/04/2020 07:59:42",
      "content": "<p>I trained a model based on Xception and MTCNN. My Notebook runs successfully on my local machine, and I guarantee that the model will run in less than nine hours. But when it comes to submission, this problem keeps coming up. Does anybody know the reason for this?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2Fe2e21cd7c43deb1b90fe9f7be54fec1b%2F1578124856714.jpg?generation=1578124879206163&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I trained a model based on Xception and MTCNN. My Notebook runs successfully on my local machine, and I guarantee that the model will run in less than nine hours. But when it comes to submission, this problem keeps coming up. Does anybody know the reason for this?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2Fe2e21cd7c43deb1b90fe9f7be54fec1b%2F1578124856714.jpg?generation=1578124879206163&amp;alt=media)",
      "votes": null
    },
    {
      "id": "710013",
      "postDate": "01/04/2020 08:12:25",
      "content": "<p>BTW, this problem occurs only when I use mtcnn, but it turns out successfully if I use dlib instead. However, dlib will take much longer to run than MTCNN.</p>",
      "rawMarkdown": "BTW, this problem occurs only when I use mtcnn, but it turns out successfully if I use dlib instead. However, dlib will take much longer to run than MTCNN.",
      "votes": null
    },
    {
      "id": "710157",
      "postDate": "01/04/2020 11:19:32",
      "content": "<p>Maybe you ran out of allowed GPU for this week?</p>",
      "rawMarkdown": "Maybe you ran out of allowed GPU for this week?",
      "votes": null
    },
    {
      "id": "710569",
      "postDate": "01/04/2020 22:35:33",
      "content": "<p>Did you step thru the kernel with success before you hit commit?  Seems like I have gotten this with I used to much RAM.</p>",
      "rawMarkdown": "Did you step thru the kernel with success before you hit commit?  Seems like I have gotten this with I used to much RAM.",
      "votes": null
    },
    {
      "id": "710658",
      "postDate": "01/05/2020 03:42:20",
      "content": "<p>My kernel runs on CPU. In addition, I have more than 20 hours of GPU quota this week.</p>",
      "rawMarkdown": "My kernel runs on CPU. In addition, I have more than 20 hours of GPU quota this week.",
      "votes": null
    },
    {
      "id": "710660",
      "postDate": "01/05/2020 03:45:54",
      "content": "<p>My kernel can run successfully step by step. This error will not occur if I change the face detection model from MTCNN to dlib. Maybe MTCNN needs more RAN than dlib.</p>",
      "rawMarkdown": "My kernel can run successfully step by step. This error will not occur if I change the face detection model from MTCNN to dlib. Maybe MTCNN needs more RAN than dlib.",
      "votes": null
    },
    {
      "id": "710693",
      "postDate": "01/05/2020 04:50:13",
      "content": "<p>OK - a bit confused by your answer.  If you step by step thru you do not get any error if dlib.</p>\n\n<p>When you step by step with MTCNN you get error?    </p>\n\n<p>You can click on the small RAM icon on the top right edge of the kernel header and it will show you the RAM usage as you step by step thru the kernel. </p>",
      "rawMarkdown": "OK - a bit confused by your answer.  If you step by step thru you do not get any error if dlib.\n\nWhen you step by step with MTCNN you get error?    \n\nYou can click on the small RAM icon on the top right edge of the kernel header and it will show you the RAM usage as you step by step thru the kernel.",
      "votes": null
    },
    {
      "id": "712475",
      "postDate": "01/07/2020 09:30:55",
      "content": "<p>I was able to run successfully step by step with MTCNN on both my local machine and the kernel on kaggle. This error occurs only during testing on Public Test Set.</p>",
      "rawMarkdown": "I was able to run successfully step by step with MTCNN on both my local machine and the kernel on kaggle. This error occurs only during testing on Public Test Set.",
      "votes": null
    },
    {
      "id": "712766",
      "postDate": "01/07/2020 15:08:00",
      "content": "<p>It is most likely that for whatever reason you have some code that occupy resources (i.e. RAM) cumulatively as your notebook goes through the 4000 hidden test set. </p>\n\n<p>I had the same error once when I was reading a 10 frames per video and feed them into two models. Ran successfully on the public test set, but came up with the same error as yours after running for 3 hours during the submission process.</p>\n\n<p>I later changed it to only read 5 frames per video and then it was completed successfully.</p>\n\n<p>I suggest to find a way to profile your resource usage on a local validation set that is sufficiently large, so that you can perform some meaningful diagnostic</p>",
      "rawMarkdown": "It is most likely that for whatever reason you have some code that occupy resources (i.e. RAM) cumulatively as your notebook goes through the 4000 hidden test set. \n\nI had the same error once when I was reading a 10 frames per video and feed them into two models. Ran successfully on the public test set, but came up with the same error as yours after running for 3 hours during the submission process.\n\nI later changed it to only read 5 frames per video and then it was completed successfully.\n\nI suggest to find a way to profile your resource usage on a local validation set that is sufficiently large, so that you can perform some meaningful diagnostic",
      "votes": null
    },
    {
      "id": "712958",
      "postDate": "01/07/2020 19:00:02",
      "content": "<p>Good suggestion - I had plans to make my local \"test\" contain 4000 to 6000 files to check for timing, etc. but thinking now that I need to drag out my old retired PC and set it up to match Kaggle restrictions.  For example, my local machines(s) never hits a RAM limit since I got a ton of RAM and did a huge swap/virtual memory in addition.  But that old piece of crap of mine only had 16GB RAM, etc.</p>\n\n<p>It's bad enough that we get no error feedback, but than to lose one of our two submissions for the day is a real kick.  </p>\n\n<p>Running my kernels on an old crap PC machine is a perfect match to kaggle limits.</p>",
      "rawMarkdown": "Good suggestion - I had plans to make my local \"test\" contain 4000 to 6000 files to check for timing, etc. but thinking now that I need to drag out my old retired PC and set it up to match Kaggle restrictions.  For example, my local machines(s) never hits a RAM limit since I got a ton of RAM and did a huge swap/virtual memory in addition.  But that old piece of crap of mine only had 16GB RAM, etc.\n\nIt's bad enough that we get no error feedback, but than to lose one of our two submissions for the day is a real kick.  \n\nRunning my kernels on an old crap PC machine is a perfect match to kaggle limits.",
      "votes": null
    },
    {
      "id": "712967",
      "postDate": "01/07/2020 19:19:42",
      "content": "<p>something like this might help? <a href=\"https://pypi.org/project/memory-profiler/\">https://pypi.org/project/memory-profiler/</a>\nWill try it out myself</p>",
      "rawMarkdown": "something like this might help? https://pypi.org/project/memory-profiler/\nWill try it out myself",
      "votes": null
    },
    {
      "id": "713108",
      "postDate": "01/07/2020 22:49:28",
      "content": "<p>Keeping record might be OK - a quick sneak over the 16 GB line might not be noticed but a complete wander over the line will work using profiler - did not look it over close enough to see the Jupyter usage.</p>\n\n<p>Think I still going with my old machine for my final test bed - going to build a PC for daughter this weekend and will have my tools spread all over the room - good time to get the old machine up and running with Ubuntu.  I don't normally use kaggle kernels because I am so bad a memory management :)</p>",
      "rawMarkdown": "Keeping record might be OK - a quick sneak over the 16 GB line might not be noticed but a complete wander over the line will work using profiler - did not look it over close enough to see the Jupyter usage.\n\nThink I still going with my old machine for my final test bed - going to build a PC for daughter this weekend and will have my tools spread all over the room - good time to get the old machine up and running with Ubuntu.  I don't normally use kaggle kernels because I am so bad a memory management :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 710013,
      "author_name": "chenshen03",
      "author_url": "",
      "post_date": "01/04/2020 08:12:25",
      "content": "<p>BTW, this problem occurs only when I use mtcnn, but it turns out successfully if I use dlib instead. However, dlib will take much longer to run than MTCNN.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 710157,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "01/04/2020 11:19:32",
      "content": "<p>Maybe you ran out of allowed GPU for this week?</p>",
      "votes": null,
      "replies": [
        {
          "id": 710658,
          "author_name": "chenshen03",
          "author_url": "",
          "post_date": "01/05/2020 03:42:20",
          "content": "<p>My kernel runs on CPU. In addition, I have more than 20 hours of GPU quota this week.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 710569,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/04/2020 22:35:33",
      "content": "<p>Did you step thru the kernel with success before you hit commit?  Seems like I have gotten this with I used to much RAM.</p>",
      "votes": null,
      "replies": [
        {
          "id": 710660,
          "author_name": "chenshen03",
          "author_url": "",
          "post_date": "01/05/2020 03:45:54",
          "content": "<p>My kernel can run successfully step by step. This error will not occur if I change the face detection model from MTCNN to dlib. Maybe MTCNN needs more RAN than dlib.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710693,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/05/2020 04:50:13",
          "content": "<p>OK - a bit confused by your answer.  If you step by step thru you do not get any error if dlib.</p>\n\n<p>When you step by step with MTCNN you get error?    </p>\n\n<p>You can click on the small RAM icon on the top right edge of the kernel header and it will show you the RAM usage as you step by step thru the kernel. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712475,
          "author_name": "chenshen03",
          "author_url": "",
          "post_date": "01/07/2020 09:30:55",
          "content": "<p>I was able to run successfully step by step with MTCNN on both my local machine and the kernel on kaggle. This error occurs only during testing on Public Test Set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712766,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "01/07/2020 15:08:00",
          "content": "<p>It is most likely that for whatever reason you have some code that occupy resources (i.e. RAM) cumulatively as your notebook goes through the 4000 hidden test set. </p>\n\n<p>I had the same error once when I was reading a 10 frames per video and feed them into two models. Ran successfully on the public test set, but came up with the same error as yours after running for 3 hours during the submission process.</p>\n\n<p>I later changed it to only read 5 frames per video and then it was completed successfully.</p>\n\n<p>I suggest to find a way to profile your resource usage on a local validation set that is sufficiently large, so that you can perform some meaningful diagnostic</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712958,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/07/2020 19:00:02",
          "content": "<p>Good suggestion - I had plans to make my local \"test\" contain 4000 to 6000 files to check for timing, etc. but thinking now that I need to drag out my old retired PC and set it up to match Kaggle restrictions.  For example, my local machines(s) never hits a RAM limit since I got a ton of RAM and did a huge swap/virtual memory in addition.  But that old piece of crap of mine only had 16GB RAM, etc.</p>\n\n<p>It's bad enough that we get no error feedback, but than to lose one of our two submissions for the day is a real kick.  </p>\n\n<p>Running my kernels on an old crap PC machine is a perfect match to kaggle limits.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 712967,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "01/07/2020 19:19:42",
          "content": "<p>something like this might help? <a href=\"https://pypi.org/project/memory-profiler/\">https://pypi.org/project/memory-profiler/</a>\nWill try it out myself</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713108,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/07/2020 22:49:28",
          "content": "<p>Keeping record might be OK - a quick sneak over the 16 GB line might not be noticed but a complete wander over the line will work using profiler - did not look it over close enough to see the Jupyter usage.</p>\n\n<p>Think I still going with my old machine for my final test bed - going to build a PC for daughter this weekend and will have my tools spread all over the room - good time to get the old machine up and running with Ubuntu.  I don't normally use kaggle kernels because I am so bad a memory management :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "710006": "I trained a model based on Xception and MTCNN. My Notebook runs successfully on my local machine, and I guarantee that the model will run in less than nine hours. But when it comes to submission, this problem keeps coming up. Does anybody know the reason for this?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2Fe2e21cd7c43deb1b90fe9f7be54fec1b%2F1578124856714.jpg?generation=1578124879206163&amp;alt=media)",
    "710013": "BTW, this problem occurs only when I use mtcnn, but it turns out successfully if I use dlib instead. However, dlib will take much longer to run than MTCNN.",
    "710157": "Maybe you ran out of allowed GPU for this week?",
    "710569": "Did you step thru the kernel with success before you hit commit?  Seems like I have gotten this with I used to much RAM.",
    "710658": "My kernel runs on CPU. In addition, I have more than 20 hours of GPU quota this week.",
    "710660": "My kernel can run successfully step by step. This error will not occur if I change the face detection model from MTCNN to dlib. Maybe MTCNN needs more RAN than dlib.",
    "710693": "OK - a bit confused by your answer.  If you step by step thru you do not get any error if dlib.\n\nWhen you step by step with MTCNN you get error?    \n\nYou can click on the small RAM icon on the top right edge of the kernel header and it will show you the RAM usage as you step by step thru the kernel.",
    "712475": "I was able to run successfully step by step with MTCNN on both my local machine and the kernel on kaggle. This error occurs only during testing on Public Test Set.",
    "712766": "It is most likely that for whatever reason you have some code that occupy resources (i.e. RAM) cumulatively as your notebook goes through the 4000 hidden test set. \n\nI had the same error once when I was reading a 10 frames per video and feed them into two models. Ran successfully on the public test set, but came up with the same error as yours after running for 3 hours during the submission process.\n\nI later changed it to only read 5 frames per video and then it was completed successfully.\n\nI suggest to find a way to profile your resource usage on a local validation set that is sufficiently large, so that you can perform some meaningful diagnostic",
    "712958": "Good suggestion - I had plans to make my local \"test\" contain 4000 to 6000 files to check for timing, etc. but thinking now that I need to drag out my old retired PC and set it up to match Kaggle restrictions.  For example, my local machines(s) never hits a RAM limit since I got a ton of RAM and did a huge swap/virtual memory in addition.  But that old piece of crap of mine only had 16GB RAM, etc.\n\nIt's bad enough that we get no error feedback, but than to lose one of our two submissions for the day is a real kick.  \n\nRunning my kernels on an old crap PC machine is a perfect match to kaggle limits.",
    "712967": "something like this might help? https://pypi.org/project/memory-profiler/\nWill try it out myself",
    "713108": "Keeping record might be OK - a quick sneak over the 16 GB line might not be noticed but a complete wander over the line will work using profiler - did not look it over close enough to see the Jupyter usage.\n\nThink I still going with my old machine for my final test bed - going to build a PC for daughter this weekend and will have my tools spread all over the room - good time to get the old machine up and running with Ubuntu.  I don't normally use kaggle kernels because I am so bad a memory management :)"
  },
  "source": "meta"
}