{
  "id": 125612,
  "title": "What does \"Kaggle Error\" means?",
  "url": "/competitions/deepfake-detection-challenge/discussion/125612",
  "author_name": "FrazierLei",
  "post_date": "2020-01-12T08:16:56.351000",
  "votes": 6,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Does it mean my program had run for more than 9 hours? \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1258264%2F5fd082c547fb14f262a040724e290d0f%2FQQ20200112160934.png?generation=1578816784560555&amp;alt=media\" alt=\"\"></p>\n\n<p>I once received this two error messages, and they had been solved.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1258264%2Fa043adf0dc2bdd42a91243d1bd5d1824%2FQQ20200112161221.png?generation=1578816782875322&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1258264%2F8269635030e75353fed3d81302919061%2FQQ20200112161213.png?generation=1578816783047634&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 716766,
      "postDate": "2020-01-12T08:16:56.350Z",
      "content": "<p>Does it mean my program had run for more than 9 hours? \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1258264%2F5fd082c547fb14f262a040724e290d0f%2FQQ20200112160934.png?generation=1578816784560555&amp;alt=media\" alt=\"\"></p>\n\n<p>I once received this two error messages, and they had been solved.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1258264%2Fa043adf0dc2bdd42a91243d1bd5d1824%2FQQ20200112161221.png?generation=1578816782875322&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1258264%2F8269635030e75353fed3d81302919061%2FQQ20200112161213.png?generation=1578816783047634&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Does it mean my program had run for more than 9 hours? \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1258264%2F5fd082c547fb14f262a040724e290d0f%2FQQ20200112160934.png?generation=1578816784560555&amp;alt=media)\n\nI once received this two error messages, and they had been solved.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1258264%2Fa043adf0dc2bdd42a91243d1bd5d1824%2FQQ20200112161221.png?generation=1578816782875322&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1258264%2F8269635030e75353fed3d81302919061%2FQQ20200112161213.png?generation=1578816783047634&amp;alt=media)\n",
      "votes": 6
    },
    {
      "id": 717287,
      "postDate": "2020-01-13T01:09:41.360Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2F61048ef3bfea675bc6e57f92011ce491%2Finbox_3117946_e2e21cd7c43deb1b90fe9f7be54fec1b_1578124856714.jpg?generation=1578877705803369&amp;alt=media\" alt=\"\"></p>\n\n<p>I have a similar error. Can someone tell me why?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2F61048ef3bfea675bc6e57f92011ce491%2Finbox_3117946_e2e21cd7c43deb1b90fe9f7be54fec1b_1578124856714.jpg?generation=1578877705803369&amp;alt=media)\n\nI have a similar error. Can someone tell me why?",
      "votes": 1,
      "replies": [
        {
          "id": 717489,
          "postDate": "2020-01-13T07:32:49.423Z",
          "content": "<p>I think you should make sure everytime you <code>to(device)</code>, the <code>device</code> should be <code>cuda:0</code>. Because we are only allowed to use one GPU.</p>",
          "rawMarkdown": "I think you should make sure everytime you `to(device)`, the `device` should be `cuda:0`. Because we are only allowed to use one GPU.",
          "votes": 1
        },
        {
          "id": 717544,
          "postDate": "2020-01-13T09:11:48.710Z",
          "content": "<p>&gt; <strong>Shen Chen wrote:</strong>\n&gt; \n&gt; <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2F61048ef3bfea675bc6e57f92011ce491%2Finbox_3117946_e2e21cd7c43deb1b90fe9f7be54fec1b_1578124856714.jpg?generation=1578877705803369&amp;alt=media\" alt=\"\">\n&gt; \n&gt; I have a similar error. Can someone tell me why?</p>\n\n<p>In my experience, this is to do with exceeding one type of resource limit that is applied on the notebook. I have two situations of the same message so far, both time I was reading 10 frames for each video to do average prediction on multiple models.  In both situations, I have managed rerun successfully, once by reducing the number of frame I read per video, and once by reducing number of models I use. I believe this is mostly related to RAM usage </p>",
          "rawMarkdown": "&gt; **Shen Chen wrote:**\n&gt; \n&gt; ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2F61048ef3bfea675bc6e57f92011ce491%2Finbox_3117946_e2e21cd7c43deb1b90fe9f7be54fec1b_1578124856714.jpg?generation=1578877705803369&amp;alt=media)\n&gt; \n&gt; I have a similar error. Can someone tell me why?\n\nIn my experience, this is to do with exceeding one type of resource limit that is applied on the notebook. I have two situations of the same message so far, both time I was reading 10 frames for each video to do average prediction on multiple models.  In both situations, I have managed rerun successfully, once by reducing the number of frame I read per video, and once by reducing number of models I use. I believe this is mostly related to RAM usage \n"
        },
        {
          "id": 717558,
          "postDate": "2020-01-13T09:30:18.040Z",
          "content": "<p><a href=\"/yifanxie\">@yifanxie</a> I only use 4 frames for each video and one model to prediction. I think these settings are unlikely to exceed the RAM usage limit. In addition, I change the face detector from mtcnn to dlib, this error will not occur, but I still couldn't find out why.</p>",
          "rawMarkdown": "@yifanxie I only use 4 frames for each video and one model to prediction. I think these settings are unlikely to exceed the RAM usage limit. In addition, I change the face detector from mtcnn to dlib, this error will not occur, but I still couldn't find out why.\n\n"
        },
        {
          "id": 717569,
          "postDate": "2020-01-13T09:49:17.200Z",
          "content": "<p>sure, everyone's setting might be different - that was just my experience so may not be directly relevant in your situation.</p>\n\n<p>I would definitely try to profile your RAM or Hardisk usage though - i.e. are you generate output file on the harddisk that might exceed the limitation? </p>\n\n<p>The message itself is quite clear in a sense that you are using too much of something when predicting on the 4000 private test set.  You can always create a routine on the public test set of 400 videos for several repetitive runs to \"simulate\" the full run, and monitor your RAM/Hardisk/CPU usage on the editing mode of the notebook. </p>",
          "rawMarkdown": "sure, everyone's setting might be different - that was just my experience so may not be directly relevant in your situation.\n\nI would definitely try to profile your RAM or Hardisk usage though - i.e. are you generate output file on the harddisk that might exceed the limitation? \n\nThe message itself is quite clear in a sense that you are using too much of something when predicting on the 4000 private test set.  You can always create a routine on the public test set of 400 videos for several repetitive runs to \"simulate\" the full run, and monitor your RAM/Hardisk/CPU usage on the editing mode of the notebook. "
        }
      ]
    },
    {
      "id": 720901,
      "postDate": "2020-01-16T20:49:23.337Z",
      "content": "<p>I also got this error. Only difference from my previous submission was the weights of my model. Really weird. </p>\n\n<p>Edit: Just to add that my notebook ran to way less than 9 hours prior to giving the \"Kaggle Error\", so not the same case as <a href=\"/carlossouza\">@carlossouza</a> described. </p>\n\n<p>Also, <a href=\"/carlossouza\">@carlossouza</a> if you are using a GPU to do face detection, I'm not sure if you are just increasing the batch size or using a fixed batch size. If you are increasing the batch size it shouldn't take 3x the time, as it doesn't scale linearly because the GPU should process the batch in parallel. The same should be true when doing the predictions. I've experimented with different numbers of faces and going from a small number to a much bigger number increased the time by 1~3 minutes in total (for the 400 videos in public test set).</p>",
      "rawMarkdown": "I also got this error. Only difference from my previous submission was the weights of my model. Really weird. \n\nEdit: Just to add that my notebook ran to way less than 9 hours prior to giving the \"Kaggle Error\", so not the same case as @carlossouza described. \n\nAlso, @carlossouza if you are using a GPU to do face detection, I'm not sure if you are just increasing the batch size or using a fixed batch size. If you are increasing the batch size it shouldn't take 3x the time, as it doesn't scale linearly because the GPU should process the batch in parallel. The same should be true when doing the predictions. I've experimented with different numbers of faces and going from a small number to a much bigger number increased the time by 1~3 minutes in total (for the 400 videos in public test set).",
      "votes": 2,
      "replies": [
        {
          "id": 721990,
          "postDate": "2020-01-18T01:02:33.690Z",
          "content": "<p>Thanks <a href=\"/pedromb\">@pedromb</a> ! You are right: it doesn't scale linearly! \nFollowing <a href=\"/juliaelliott\">@juliaelliott</a> 's advice, I re-run it and now it worked.\nSo, my first conclusion was incorrect: \"Kaggle Error\" does NOT mean we blow up 9hs time.</p>",
          "rawMarkdown": "Thanks @pedromb ! You are right: it doesn't scale linearly! \nFollowing @juliaelliott 's advice, I re-run it and now it worked.\nSo, my first conclusion was incorrect: \"Kaggle Error\" does NOT mean we blow up 9hs time."
        }
      ]
    },
    {
      "id": 719903,
      "postDate": "2020-01-16T01:00:16.050Z",
      "content": "<p>“Kaggle Error” is a rare system error, and it’s recommended that you resubmit to resolve the issue. Please let us know if resubmission continues to fail.</p>",
      "rawMarkdown": "“Kaggle Error” is a rare system error, and it’s recommended that you resubmit to resolve the issue. Please let us know if resubmission continues to fail.",
      "votes": 2,
      "replies": [
        {
          "id": 721989,
          "postDate": "2020-01-18T01:00:50.970Z",
          "content": "<p>Thank you <a href=\"/juliaelliott\">@juliaelliott</a> ! Thanks to your message, I re-run the notebook, and now it worked!</p>",
          "rawMarkdown": "Thank you @juliaelliott ! Thanks to your message, I re-run the notebook, and now it worked!"
        },
        {
          "id": 722556,
          "postDate": "2020-01-18T17:54:05.993Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> It looks <a href=\"/carlossouza\">@carlossouza</a> is not alone to have this error today. Same kernel run twice today, first attempt with \"Kaggle error\" after a few hours, second attempt \"Notebook Exceeded Allowed compute\" after a few minutes.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2Fcf3c0b5933cae3b49a96025a25a00d77%2Fkaggle_error.png?generation=1579369964893669&amp;alt=media\" alt=\"\"></p>\n\n<p>Quite difficult to troubleshoot.</p>",
          "rawMarkdown": "@juliaelliott It looks @carlossouza is not alone to have this error today. Same kernel run twice today, first attempt with \"Kaggle error\" after a few hours, second attempt \"Notebook Exceeded Allowed compute\" after a few minutes.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2Fcf3c0b5933cae3b49a96025a25a00d77%2Fkaggle_error.png?generation=1579369964893669&amp;alt=media)\n\nQuite difficult to troubleshoot.",
          "votes": 1
        },
        {
          "id": 722570,
          "postDate": "2020-01-18T18:28:27Z",
          "content": "<p>Very difficult to troubleshoot. My last 10(!) submissions are errors, I am trying different things, but with 2 submissions a day it is proving too hard to pinpoint the problem. </p>\n\n<p>It can be super nice if someone from Kaggle looks at 10-20 of the common failures and write a short guide about the pitfalls that we need to be mindful about. </p>",
          "rawMarkdown": "Very difficult to troubleshoot. My last 10(!) submissions are errors, I am trying different things, but with 2 submissions a day it is proving too hard to pinpoint the problem. \n\nIt can be super nice if someone from Kaggle looks at 10-20 of the common failures and write a short guide about the pitfalls that we need to be mindful about. ",
          "votes": 3
        },
        {
          "id": 722582,
          "postDate": "2020-01-18T18:55:52.917Z",
          "content": "<p><a href=\"/zaharch\">@zaharch</a> To predict safely, I've surrounded <code>model.predict</code> for each video by <code>try/except</code>, I've added <code>np.clip</code> to avoid log loss computation issue but it looks not enough. I think most of my problems comes batch_size, frames buffer and face size, but I cannot understand why it works on commit. Maybe some faces extracted are too big (only on test dataset loaded on submit) that makes memory crash (either RAM and/or GPU). </p>\n\n<p><code>\nfor idx, row in tqdm(submission_pd.iterrows(), total=submission_pd.shape[0]):\n    filename = row[\"filename\"]\n    try: <br>\n        # Extract faces from video frames and prepare for model (normalize)\n        ds_test= VideoDataset(filename, label=None, subset='test', transform=tr) <br>\n        # Run model\n        dl_test = DataLoader(ds_test, batch_size=24, shuffle=False, num_workers=1) <br>\n        clf_predict_probas, _ = run_model_fn(dl_test ) <br>\n        ds_test.cleanup()\n        # Average mean (of all faces within all frames, other strategy to try out)\n        filename_prob = np.nanmean(clf_predict_probas)\n        submission_pd.loc[idx, \"label\"] = np.clip(filename_prob, 0.01, 0.99)\n        torch.cuda.empty_cache() # If any GPU issue\n    except Exception as ex:\n        submission_pd.loc[idx, \"label\"] = 0.5\n        print(\"Cannot predict on %s\" % filename) \n</code>\nAnd <code>np.fillna</code> if any prediction totally fails (no error but no face or something else). </p>\n\n<p><code>\nsubmission_pd.fillna(0.5, inplace=True)\nsubmission_pd.sort_values('filename').to_csv('submission.csv', index=False)\n</code></p>\n\n<p><a href=\"/juliaelliott\">@juliaelliott</a> I've read in another thread that we should read <code>test_videos</code> folder instead of <code>sample_submission.csv</code>, is it true? I would expect getting filenames from <code>sample_submission.csv</code> like in other competitions.</p>\n\n<p>```\nTEST_HOME = \"/kaggle/input/deepfake-detection-challenge/test_videos/\"\nSAMPLE_SUBMISSION = \"/kaggle/input/deepfake-detection-challenge/sample_submission.csv\"</p>\n\n<h1>I would expect this one to work</h1>\n\n<h1>submission_pd = pd.read_csv(SAMPLE_SUBMISSION)</h1>\n\n<h1>But this is what I do:</h1>\n\n<p>filenames = glob.glob(TEST_HOME + \"*.mp4\")\nbasenames = [os.path.basename(filename) for filename in filenames]\nsubmission_pd = pd.DataFrame(basenames, columns=[\"filename\"])\n```</p>",
          "rawMarkdown": "@zaharch To predict safely, I've surrounded `model.predict` for each video by `try/except`, I've added `np.clip` to avoid log loss computation issue but it looks not enough. I think most of my problems comes batch_size, frames buffer and face size, but I cannot understand why it works on commit. Maybe some faces extracted are too big (only on test dataset loaded on submit) that makes memory crash (either RAM and/or GPU). \n\n```\nfor idx, row in tqdm(submission_pd.iterrows(), total=submission_pd.shape[0]):\n\tfilename = row[\"filename\"]\n\ttry:                \n\t\t# Extract faces from video frames and prepare for model (normalize)\n\t\tds_test= VideoDataset(filename, label=None, subset='test', transform=tr)                \n\t\t# Run model\n\t\tdl_test = DataLoader(ds_test, batch_size=24, shuffle=False, num_workers=1)        \n\t\tclf_predict_probas, _ = run_model_fn(dl_test )                \n\t\tds_test.cleanup()\n\t\t# Average mean (of all faces within all frames, other strategy to try out)\n\t\tfilename_prob = np.nanmean(clf_predict_probas)\n\t\tsubmission_pd.loc[idx, \"label\"] = np.clip(filename_prob, 0.01, 0.99)\n\t\ttorch.cuda.empty_cache() # If any GPU issue\n\texcept Exception as ex:\n\t\tsubmission_pd.loc[idx, \"label\"] = 0.5\n\t\tprint(\"Cannot predict on %s\" % filename) \n```\nAnd `np.fillna` if any prediction totally fails (no error but no face or something else). \n\n```\nsubmission_pd.fillna(0.5, inplace=True)\nsubmission_pd.sort_values('filename').to_csv('submission.csv', index=False)\n```\n\n@juliaelliott I've read in another thread that we should read `test_videos` folder instead of `sample_submission.csv`, is it true? I would expect getting filenames from `sample_submission.csv` like in other competitions.\n\n```\nTEST_HOME = \"/kaggle/input/deepfake-detection-challenge/test_videos/\"\nSAMPLE_SUBMISSION = \"/kaggle/input/deepfake-detection-challenge/sample_submission.csv\"\n\n# I would expect this one to work\n# submission_pd = pd.read_csv(SAMPLE_SUBMISSION) \n\n# But this is what I do:\nfilenames = glob.glob(TEST_HOME + \"*.mp4\")\nbasenames = [os.path.basename(filename) for filename in filenames]\nsubmission_pd = pd.DataFrame(basenames, columns=[\"filename\"])\n```\n\n",
          "votes": 2
        },
        {
          "id": 722728,
          "postDate": "2020-01-19T02:11:08.503Z",
          "content": "<p>Thank you, this is extremely helpful and should be part of the competition guidelines. I think the latter way is the right way because you're not dependent on a submission file with a priori labels. Note that this question has been raised in other discussion topics without followup from the organizers.</p>",
          "rawMarkdown": "Thank you, this is extremely helpful and should be part of the competition guidelines. I think the latter way is the right way because you're not dependent on a submission file with a priori labels. Note that this question has been raised in other discussion topics without followup from the organizers."
        },
        {
          "id": 767973,
          "postDate": "2020-03-10T10:13:48.487Z",
          "content": "<p>Hi, I have encountered the kaggle error twice with only ckpt changed. It wasted two times of submissions. Can you help me out with this? <a href=\"/juliaelliott\">@juliaelliott</a> </p>",
          "rawMarkdown": "Hi, I have encountered the kaggle error twice with only ckpt changed. It wasted two times of submissions. Can you help me out with this? @juliaelliott "
        }
      ]
    },
    {
      "id": 717042,
      "postDate": "2020-01-12T16:36:12.710Z",
      "content": "<p>I just received the 1st one (Kaggle Error). The only thing I've changed vs. a previously successful submission was the # of faces extracted from the test videos (3x vs. my baseline), to make the inference by averaging. I was actually expecting an error, because my baseline was taking 4hs to complete; therefore, 3x-ing the time to extract faces from videos should approx. 3x the time to complete, i.e. 12hs, above the 9hs threshold. </p>\n\n<p>This experiment confirmed 2 things for me:\n1. This \"Kaggle Error\" message is what the system outputs when we blow up 9hs\n2. I need a faster Face Detector to make inferences using more frames/video :)</p>\n\n<p>Hope it helps! Cheers! :)</p>",
      "rawMarkdown": "I just received the 1st one (Kaggle Error). The only thing I've changed vs. a previously successful submission was the # of faces extracted from the test videos (3x vs. my baseline), to make the inference by averaging. I was actually expecting an error, because my baseline was taking 4hs to complete; therefore, 3x-ing the time to extract faces from videos should approx. 3x the time to complete, i.e. 12hs, above the 9hs threshold. \n\nThis experiment confirmed 2 things for me:\n1. This \"Kaggle Error\" message is what the system outputs when we blow up 9hs\n2. I need a faster Face Detector to make inferences using more frames/video :)\n\nHope it helps! Cheers! :)",
      "votes": 2,
      "replies": [
        {
          "id": 717490,
          "postDate": "2020-01-13T07:33:41.497Z",
          "content": "<p>Thanks! It does make sense.</p>",
          "rawMarkdown": "Thanks! It does make sense.",
          "votes": 2
        }
      ]
    },
    {
      "id": 736629,
      "postDate": "2020-02-04T12:02:52.673Z",
      "content": "<p>I also got \"Kaggle Error\" now. I'll try to resubmit...</p>",
      "rawMarkdown": "I also got \"Kaggle Error\" now. I'll try to resubmit..."
    },
    {
      "id": 717590,
      "postDate": "2020-01-13T10:47:06.153Z",
      "content": "<p>The Kaggle Error is the error on your training set usually</p>",
      "rawMarkdown": "The Kaggle Error is the error on your training set usually"
    },
    {
      "id": 722726,
      "postDate": "2020-01-19T02:02:40.923Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 717287,
      "author_name": "Chason",
      "author_url": "",
      "post_date": "2020-01-13T01:09:41.360000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2F61048ef3bfea675bc6e57f92011ce491%2Finbox_3117946_e2e21cd7c43deb1b90fe9f7be54fec1b_1578124856714.jpg?generation=1578877705803369&amp;alt=media\" alt=\"\"></p>\n\n<p>I have a similar error. Can someone tell me why?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 717489,
          "author_name": "FrazierLei",
          "author_url": "",
          "post_date": "2020-01-13T07:32:49.423000",
          "content": "<p>I think you should make sure everytime you <code>to(device)</code>, the <code>device</code> should be <code>cuda:0</code>. Because we are only allowed to use one GPU.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 717544,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-13T09:11:48.710000",
          "content": "<p>&gt; <strong>Shen Chen wrote:</strong>\n&gt; \n&gt; <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2F61048ef3bfea675bc6e57f92011ce491%2Finbox_3117946_e2e21cd7c43deb1b90fe9f7be54fec1b_1578124856714.jpg?generation=1578877705803369&amp;alt=media\" alt=\"\">\n&gt; \n&gt; I have a similar error. Can someone tell me why?</p>\n\n<p>In my experience, this is to do with exceeding one type of resource limit that is applied on the notebook. I have two situations of the same message so far, both time I was reading 10 frames for each video to do average prediction on multiple models.  In both situations, I have managed rerun successfully, once by reducing the number of frame I read per video, and once by reducing number of models I use. I believe this is mostly related to RAM usage </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 717558,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-01-13T09:30:18.040000",
          "content": "<p><a href=\"/yifanxie\">@yifanxie</a> I only use 4 frames for each video and one model to prediction. I think these settings are unlikely to exceed the RAM usage limit. In addition, I change the face detector from mtcnn to dlib, this error will not occur, but I still couldn't find out why.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 717569,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-13T09:49:17.200000",
          "content": "<p>sure, everyone's setting might be different - that was just my experience so may not be directly relevant in your situation.</p>\n\n<p>I would definitely try to profile your RAM or Hardisk usage though - i.e. are you generate output file on the harddisk that might exceed the limitation? </p>\n\n<p>The message itself is quite clear in a sense that you are using too much of something when predicting on the 4000 private test set.  You can always create a routine on the public test set of 400 videos for several repetitive runs to \"simulate\" the full run, and monitor your RAM/Hardisk/CPU usage on the editing mode of the notebook. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 720901,
      "author_name": "Pedro Bernardo",
      "author_url": "",
      "post_date": "2020-01-16T20:49:23.337000",
      "content": "<p>I also got this error. Only difference from my previous submission was the weights of my model. Really weird. </p>\n\n<p>Edit: Just to add that my notebook ran to way less than 9 hours prior to giving the \"Kaggle Error\", so not the same case as <a href=\"/carlossouza\">@carlossouza</a> described. </p>\n\n<p>Also, <a href=\"/carlossouza\">@carlossouza</a> if you are using a GPU to do face detection, I'm not sure if you are just increasing the batch size or using a fixed batch size. If you are increasing the batch size it shouldn't take 3x the time, as it doesn't scale linearly because the GPU should process the batch in parallel. The same should be true when doing the predictions. I've experimented with different numbers of faces and going from a small number to a much bigger number increased the time by 1~3 minutes in total (for the 400 videos in public test set).</p>",
      "votes": 2,
      "replies": [
        {
          "id": 721990,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-18T01:02:33.690000",
          "content": "<p>Thanks <a href=\"/pedromb\">@pedromb</a> ! You are right: it doesn't scale linearly! \nFollowing <a href=\"/juliaelliott\">@juliaelliott</a> 's advice, I re-run it and now it worked.\nSo, my first conclusion was incorrect: \"Kaggle Error\" does NOT mean we blow up 9hs time.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 719903,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2020-01-16T01:00:16.050000",
      "content": "<p>“Kaggle Error” is a rare system error, and it’s recommended that you resubmit to resolve the issue. Please let us know if resubmission continues to fail.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 721989,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-18T01:00:50.970000",
          "content": "<p>Thank you <a href=\"/juliaelliott\">@juliaelliott</a> ! Thanks to your message, I re-run the notebook, and now it worked!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 722556,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-01-18T17:54:05.993000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> It looks <a href=\"/carlossouza\">@carlossouza</a> is not alone to have this error today. Same kernel run twice today, first attempt with \"Kaggle error\" after a few hours, second attempt \"Notebook Exceeded Allowed compute\" after a few minutes.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2Fcf3c0b5933cae3b49a96025a25a00d77%2Fkaggle_error.png?generation=1579369964893669&amp;alt=media\" alt=\"\"></p>\n\n<p>Quite difficult to troubleshoot.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 722570,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-01-18T18:28:27",
          "content": "<p>Very difficult to troubleshoot. My last 10(!) submissions are errors, I am trying different things, but with 2 submissions a day it is proving too hard to pinpoint the problem. </p>\n\n<p>It can be super nice if someone from Kaggle looks at 10-20 of the common failures and write a short guide about the pitfalls that we need to be mindful about. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 722582,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-01-18T18:55:52.917000",
          "content": "<p><a href=\"/zaharch\">@zaharch</a> To predict safely, I've surrounded <code>model.predict</code> for each video by <code>try/except</code>, I've added <code>np.clip</code> to avoid log loss computation issue but it looks not enough. I think most of my problems comes batch_size, frames buffer and face size, but I cannot understand why it works on commit. Maybe some faces extracted are too big (only on test dataset loaded on submit) that makes memory crash (either RAM and/or GPU). </p>\n\n<p><code>\nfor idx, row in tqdm(submission_pd.iterrows(), total=submission_pd.shape[0]):\n    filename = row[\"filename\"]\n    try: <br>\n        # Extract faces from video frames and prepare for model (normalize)\n        ds_test= VideoDataset(filename, label=None, subset='test', transform=tr) <br>\n        # Run model\n        dl_test = DataLoader(ds_test, batch_size=24, shuffle=False, num_workers=1) <br>\n        clf_predict_probas, _ = run_model_fn(dl_test ) <br>\n        ds_test.cleanup()\n        # Average mean (of all faces within all frames, other strategy to try out)\n        filename_prob = np.nanmean(clf_predict_probas)\n        submission_pd.loc[idx, \"label\"] = np.clip(filename_prob, 0.01, 0.99)\n        torch.cuda.empty_cache() # If any GPU issue\n    except Exception as ex:\n        submission_pd.loc[idx, \"label\"] = 0.5\n        print(\"Cannot predict on %s\" % filename) \n</code>\nAnd <code>np.fillna</code> if any prediction totally fails (no error but no face or something else). </p>\n\n<p><code>\nsubmission_pd.fillna(0.5, inplace=True)\nsubmission_pd.sort_values('filename').to_csv('submission.csv', index=False)\n</code></p>\n\n<p><a href=\"/juliaelliott\">@juliaelliott</a> I've read in another thread that we should read <code>test_videos</code> folder instead of <code>sample_submission.csv</code>, is it true? I would expect getting filenames from <code>sample_submission.csv</code> like in other competitions.</p>\n\n<p>```\nTEST_HOME = \"/kaggle/input/deepfake-detection-challenge/test_videos/\"\nSAMPLE_SUBMISSION = \"/kaggle/input/deepfake-detection-challenge/sample_submission.csv\"</p>\n\n<h1>I would expect this one to work</h1>\n\n<h1>submission_pd = pd.read_csv(SAMPLE_SUBMISSION)</h1>\n\n<h1>But this is what I do:</h1>\n\n<p>filenames = glob.glob(TEST_HOME + \"*.mp4\")\nbasenames = [os.path.basename(filename) for filename in filenames]\nsubmission_pd = pd.DataFrame(basenames, columns=[\"filename\"])\n```</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 722728,
          "author_name": "Mark Conway",
          "author_url": "",
          "post_date": "2020-01-19T02:11:08.503000",
          "content": "<p>Thank you, this is extremely helpful and should be part of the competition guidelines. I think the latter way is the right way because you're not dependent on a submission file with a priori labels. Note that this question has been raised in other discussion topics without followup from the organizers.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767973,
          "author_name": "xxn-xx",
          "author_url": "",
          "post_date": "2020-03-10T10:13:48.487000",
          "content": "<p>Hi, I have encountered the kaggle error twice with only ckpt changed. It wasted two times of submissions. Can you help me out with this? <a href=\"/juliaelliott\">@juliaelliott</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 717042,
      "author_name": "Carlos Souza",
      "author_url": "",
      "post_date": "2020-01-12T16:36:12.710000",
      "content": "<p>I just received the 1st one (Kaggle Error). The only thing I've changed vs. a previously successful submission was the # of faces extracted from the test videos (3x vs. my baseline), to make the inference by averaging. I was actually expecting an error, because my baseline was taking 4hs to complete; therefore, 3x-ing the time to extract faces from videos should approx. 3x the time to complete, i.e. 12hs, above the 9hs threshold. </p>\n\n<p>This experiment confirmed 2 things for me:\n1. This \"Kaggle Error\" message is what the system outputs when we blow up 9hs\n2. I need a faster Face Detector to make inferences using more frames/video :)</p>\n\n<p>Hope it helps! Cheers! :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 717490,
          "author_name": "FrazierLei",
          "author_url": "",
          "post_date": "2020-01-13T07:33:41.497000",
          "content": "<p>Thanks! It does make sense.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 736629,
      "author_name": "Felipe Loque",
      "author_url": "",
      "post_date": "2020-02-04T12:02:52.673000",
      "content": "<p>I also got \"Kaggle Error\" now. I'll try to resubmit...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 717590,
      "author_name": "Oleg Gribanov",
      "author_url": "",
      "post_date": "2020-01-13T10:47:06.153000",
      "content": "<p>The Kaggle Error is the error on your training set usually</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 722726,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-19T02:02:40.923000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "716766": "Does it mean my program had run for more than 9 hours? \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1258264%2F5fd082c547fb14f262a040724e290d0f%2FQQ20200112160934.png?generation=1578816784560555&amp;alt=media)\n\nI once received this two error messages, and they had been solved.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1258264%2Fa043adf0dc2bdd42a91243d1bd5d1824%2FQQ20200112161221.png?generation=1578816782875322&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1258264%2F8269635030e75353fed3d81302919061%2FQQ20200112161213.png?generation=1578816783047634&amp;alt=media)\n",
    "717287": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2F61048ef3bfea675bc6e57f92011ce491%2Finbox_3117946_e2e21cd7c43deb1b90fe9f7be54fec1b_1578124856714.jpg?generation=1578877705803369&amp;alt=media)\n\nI have a similar error. Can someone tell me why?",
    "720901": "I also got this error. Only difference from my previous submission was the weights of my model. Really weird. \n\nEdit: Just to add that my notebook ran to way less than 9 hours prior to giving the \"Kaggle Error\", so not the same case as @carlossouza described. \n\nAlso, @carlossouza if you are using a GPU to do face detection, I'm not sure if you are just increasing the batch size or using a fixed batch size. If you are increasing the batch size it shouldn't take 3x the time, as it doesn't scale linearly because the GPU should process the batch in parallel. The same should be true when doing the predictions. I've experimented with different numbers of faces and going from a small number to a much bigger number increased the time by 1~3 minutes in total (for the 400 videos in public test set).",
    "719903": "“Kaggle Error” is a rare system error, and it’s recommended that you resubmit to resolve the issue. Please let us know if resubmission continues to fail.",
    "717042": "I just received the 1st one (Kaggle Error). The only thing I've changed vs. a previously successful submission was the # of faces extracted from the test videos (3x vs. my baseline), to make the inference by averaging. I was actually expecting an error, because my baseline was taking 4hs to complete; therefore, 3x-ing the time to extract faces from videos should approx. 3x the time to complete, i.e. 12hs, above the 9hs threshold. \n\nThis experiment confirmed 2 things for me:\n1. This \"Kaggle Error\" message is what the system outputs when we blow up 9hs\n2. I need a faster Face Detector to make inferences using more frames/video :)\n\nHope it helps! Cheers! :)",
    "736629": "I also got \"Kaggle Error\" now. I'll try to resubmit...",
    "717590": "The Kaggle Error is the error on your training set usually",
    "722726": ""
  }
}