{
  "id": 127824,
  "title": "Submission Errors Solutions",
  "url": "/competitions/deepfake-detection-challenge/discussion/127824",
  "author_name": "",
  "post_date": "2020-01-27T01:44:44.743280200Z",
  "votes": 26,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Here are some solutions to submission scoring errors, LB and CV too much difference, finishes too fast or other similar submission errors.\n1. Don't use 'sample_submission.csv', infer over test_videos dir.\n2. Make sure add try except around proccess image.\n3. Set values to all 0.5 to make sure we won't miss any file names\n4. Clean up your working dir afterwards, there's a limitation of output files.\n5. Last but not least, make sure your submission file is 'submission.csv' and add index=False parameter\nHere is an example code:</p>\n\n<p><code>\nfilenames=[]\ntest_X=[]\nfor path in os.listidr('../input/deepfake-detection-challenge/test_videos/'):\n    try:\n        processed=process_img(path)\n    except Exception as err:\n        print(err)\n        continue\n    test_X.append(processed)\n    filenames.append(path)\npreds=[]\nbatch_size=16\nfor x in range(len(test_X)//batch_size):\n    batch_X=test_X[batch_size*x:batch_size*(x+1)]\n    try:\n         pred=model.predict_on_batch(batch_X)\n    except Exception as err:\n         print(err)\n         pred=[0.5 for _ in range(batch_size)]\n    preds+=pred\nbatch_X=test_X[batch_size*(len(test_X)//16):]\ntry:\n    pred=model.predict_on_batch(batch_X)\nexcept Exception as err:\n    print(err)\n    pred=[0.5 for _ in range(batch_size)]\npreds+=pred\nsub=pd.DataFrame()\nsub['filename']=os.listidr('../input/deepfake-detection-challenge/test_videos/')\nsub['label']=0.5\nfor pred,filename in zip(preds,filenames):\n    sub.loc[sub['filename']==filename,'label']=pred\nsub.to_csv('submission.csv',index=False)\n</code></p>\n\n<p>You can comment below if you are still experiencing those kind of errors after these fixes. Be sure to provide key parts of code.</p>\n\n<p>If your error is timeout, do remember that the unseen validation set is 4000 videos, which is 10 times more. Your notebook should finish within 54 minutes(if we consider it as a linear relationship) when committing in order to finish predicting the unseen validation set. If not, be sure to turn on GPU(if that helps) and optimize your code to make it faster.</p>",
  "messages": [
    {
      "id": "730024",
      "postDate": "01/27/2020 01:44:44",
      "content": "<p>Here are some solutions to submission scoring errors, LB and CV too much difference, finishes too fast or other similar submission errors.\n1. Don't use 'sample_submission.csv', infer over test_videos dir.\n2. Make sure add try except around proccess image.\n3. Set values to all 0.5 to make sure we won't miss any file names\n4. Clean up your working dir afterwards, there's a limitation of output files.\n5. Last but not least, make sure your submission file is 'submission.csv' and add index=False parameter\nHere is an example code:</p>\n\n<p><code>\nfilenames=[]\ntest_X=[]\nfor path in os.listidr('../input/deepfake-detection-challenge/test_videos/'):\n    try:\n        processed=process_img(path)\n    except Exception as err:\n        print(err)\n        continue\n    test_X.append(processed)\n    filenames.append(path)\npreds=[]\nbatch_size=16\nfor x in range(len(test_X)//batch_size):\n    batch_X=test_X[batch_size*x:batch_size*(x+1)]\n    try:\n         pred=model.predict_on_batch(batch_X)\n    except Exception as err:\n         print(err)\n         pred=[0.5 for _ in range(batch_size)]\n    preds+=pred\nbatch_X=test_X[batch_size*(len(test_X)//16):]\ntry:\n    pred=model.predict_on_batch(batch_X)\nexcept Exception as err:\n    print(err)\n    pred=[0.5 for _ in range(batch_size)]\npreds+=pred\nsub=pd.DataFrame()\nsub['filename']=os.listidr('../input/deepfake-detection-challenge/test_videos/')\nsub['label']=0.5\nfor pred,filename in zip(preds,filenames):\n    sub.loc[sub['filename']==filename,'label']=pred\nsub.to_csv('submission.csv',index=False)\n</code></p>\n\n<p>You can comment below if you are still experiencing those kind of errors after these fixes. Be sure to provide key parts of code.</p>\n\n<p>If your error is timeout, do remember that the unseen validation set is 4000 videos, which is 10 times more. Your notebook should finish within 54 minutes(if we consider it as a linear relationship) when committing in order to finish predicting the unseen validation set. If not, be sure to turn on GPU(if that helps) and optimize your code to make it faster.</p>",
      "rawMarkdown": "Here are some solutions to submission scoring errors, LB and CV too much difference, finishes too fast or other similar submission errors.\n1. Don't use 'sample_submission.csv', infer over test_videos dir.\n2. Make sure add try except around proccess image.\n3. Set values to all 0.5 to make sure we won't miss any file names\n4. Clean up your working dir afterwards, there's a limitation of output files.\n5. Last but not least, make sure your submission file is 'submission.csv' and add index=False parameter\nHere is an example code:\n\n\n```\nfilenames=[]\ntest_X=[]\nfor path in os.listidr('../input/deepfake-detection-challenge/test_videos/'):\n    try:\n        processed=process_img(path)\n    except Exception as err:\n        print(err)\n        continue\n    test_X.append(processed)\n    filenames.append(path)\npreds=[]\nbatch_size=16\nfor x in range(len(test_X)//batch_size):\n    batch_X=test_X[batch_size*x:batch_size*(x+1)]\n    try:\n         pred=model.predict_on_batch(batch_X)\n    except Exception as err:\n         print(err)\n         pred=[0.5 for _ in range(batch_size)]\n    preds+=pred\nbatch_X=test_X[batch_size*(len(test_X)//16):]\ntry:\n    pred=model.predict_on_batch(batch_X)\nexcept Exception as err:\n    print(err)\n    pred=[0.5 for _ in range(batch_size)]\npreds+=pred\nsub=pd.DataFrame()\nsub['filename']=os.listidr('../input/deepfake-detection-challenge/test_videos/')\nsub['label']=0.5\nfor pred,filename in zip(preds,filenames):\n    sub.loc[sub['filename']==filename,'label']=pred\nsub.to_csv('submission.csv',index=False)\n```\n\nYou can comment below if you are still experiencing those kind of errors after these fixes. Be sure to provide key parts of code.\n\nIf your error is timeout, do remember that the unseen validation set is 4000 videos, which is 10 times more. Your notebook should finish within 54 minutes(if we consider it as a linear relationship) when committing in order to finish predicting the unseen validation set. If not, be sure to turn on GPU(if that helps) and optimize your code to make it faster.",
      "votes": null
    },
    {
      "id": "730066",
      "postDate": "01/27/2020 03:54:13",
      "content": "<p>You might also want to add error handling around the inference part of your code. For testing error handling on my pipeline I purposely corrupted one of the files to force it throw an error.</p>",
      "rawMarkdown": "You might also want to add error handling around the inference part of your code. For testing error handling on my pipeline I purposely corrupted one of the files to force it throw an error.",
      "votes": null
    },
    {
      "id": "730070",
      "postDate": "01/27/2020 04:05:52",
      "content": "<p>ok thanks. But even if we added error handling, how will we be able to predict the rest. So the only thing we could do is to avoid  corrupted ones when preprocessing.</p>",
      "rawMarkdown": "ok thanks. But even if we added error handling, how will we be able to predict the rest. So the only thing we could do is to avoid  corrupted ones when preprocessing.",
      "votes": null
    },
    {
      "id": "730083",
      "postDate": "01/27/2020 05:01:34",
      "content": "<p>If your using ffmpeg and subprocess calls a simple try/except will not catch corrupted file errors.</p>\n\n<p>subprocess returns a value of 0 when the call is successful - when it fails other values are returned - so I used</p>\n\n<p>a = subprocess.....</p>\n\n<p>if a ==0:\n    happyness\nelse:\n    sad - got here due to the error</p>",
      "rawMarkdown": "If your using ffmpeg and subprocess calls a simple try/except will not catch corrupted file errors.\n\nsubprocess returns a value of 0 when the call is successful - when it fails other values are returned - so I used\n\na = subprocess.....\n\nif a ==0:\n    happyness\nelse:\n    sad - got here due to the error",
      "votes": null
    },
    {
      "id": "730484",
      "postDate": "01/27/2020 14:41:45",
      "content": "<p>To clarify I mean add error handling when running inference on each batch, if there is an error running inference on the batch set results to 0.5 and continue processing the rest of the batches. </p>",
      "rawMarkdown": "To clarify I mean add error handling when running inference on each batch, if there is an error running inference on the batch set results to 0.5 and continue processing the rest of the batches.",
      "votes": null
    },
    {
      "id": "730524",
      "postDate": "01/27/2020 15:20:23",
      "content": "<p>Thanks. Added to the post.</p>",
      "rawMarkdown": "Thanks. Added to the post.",
      "votes": null
    },
    {
      "id": "737869",
      "postDate": "02/05/2020 21:17:46",
      "content": "<p>How is that even possible that the sample submission doesn't contain all files? Why this mess? Miss file names? What game is that? Maybe we don't want to play it..\n... Supposed to be data competition... </p>",
      "rawMarkdown": "How is that even possible that the sample submission doesn't contain all files? Why this mess? Miss file names? What game is that? Maybe we don't want to play it..\n... Supposed to be data competition...",
      "votes": null
    },
    {
      "id": "782180",
      "postDate": "03/22/2020 02:19:06",
      "content": "<p>Hi, I tried your error solution and the my submission finished in 1 hour, getting 0.68828, does this mean that I submitted all 0.5s?</p>",
      "rawMarkdown": "Hi, I tried your error solution and the my submission finished in 1 hour, getting 0.68828, does this mean that I submitted all 0.5s?",
      "votes": null
    },
    {
      "id": "782215",
      "postDate": "03/22/2020 03:43:27",
      "content": "<p>If the submission fails, the submission usually finishes in less than 30 minutes.\nCan you share the run time of your commit too.</p>",
      "rawMarkdown": "If the submission fails, the submission usually finishes in less than 30 minutes.\nCan you share the run time of your commit too.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 730066,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "01/27/2020 03:54:13",
      "content": "<p>You might also want to add error handling around the inference part of your code. For testing error handling on my pipeline I purposely corrupted one of the files to force it throw an error.</p>",
      "votes": null,
      "replies": [
        {
          "id": 730070,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "01/27/2020 04:05:52",
          "content": "<p>ok thanks. But even if we added error handling, how will we be able to predict the rest. So the only thing we could do is to avoid  corrupted ones when preprocessing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 730484,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "01/27/2020 14:41:45",
          "content": "<p>To clarify I mean add error handling when running inference on each batch, if there is an error running inference on the batch set results to 0.5 and continue processing the rest of the batches. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 730524,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "01/27/2020 15:20:23",
          "content": "<p>Thanks. Added to the post.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 737869,
          "author_name": "simoninparis",
          "author_url": "",
          "post_date": "02/05/2020 21:17:46",
          "content": "<p>How is that even possible that the sample submission doesn't contain all files? Why this mess? Miss file names? What game is that? Maybe we don't want to play it..\n... Supposed to be data competition... </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 730083,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/27/2020 05:01:34",
      "content": "<p>If your using ffmpeg and subprocess calls a simple try/except will not catch corrupted file errors.</p>\n\n<p>subprocess returns a value of 0 when the call is successful - when it fails other values are returned - so I used</p>\n\n<p>a = subprocess.....</p>\n\n<p>if a ==0:\n    happyness\nelse:\n    sad - got here due to the error</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 782180,
      "author_name": "",
      "author_url": "",
      "post_date": "03/22/2020 02:19:06",
      "content": "<p>Hi, I tried your error solution and the my submission finished in 1 hour, getting 0.68828, does this mean that I submitted all 0.5s?</p>",
      "votes": null,
      "replies": [
        {
          "id": 782215,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "03/22/2020 03:43:27",
          "content": "<p>If the submission fails, the submission usually finishes in less than 30 minutes.\nCan you share the run time of your commit too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "730024": "Here are some solutions to submission scoring errors, LB and CV too much difference, finishes too fast or other similar submission errors.\n1. Don't use 'sample_submission.csv', infer over test_videos dir.\n2. Make sure add try except around proccess image.\n3. Set values to all 0.5 to make sure we won't miss any file names\n4. Clean up your working dir afterwards, there's a limitation of output files.\n5. Last but not least, make sure your submission file is 'submission.csv' and add index=False parameter\nHere is an example code:\n\n\n```\nfilenames=[]\ntest_X=[]\nfor path in os.listidr('../input/deepfake-detection-challenge/test_videos/'):\n    try:\n        processed=process_img(path)\n    except Exception as err:\n        print(err)\n        continue\n    test_X.append(processed)\n    filenames.append(path)\npreds=[]\nbatch_size=16\nfor x in range(len(test_X)//batch_size):\n    batch_X=test_X[batch_size*x:batch_size*(x+1)]\n    try:\n         pred=model.predict_on_batch(batch_X)\n    except Exception as err:\n         print(err)\n         pred=[0.5 for _ in range(batch_size)]\n    preds+=pred\nbatch_X=test_X[batch_size*(len(test_X)//16):]\ntry:\n    pred=model.predict_on_batch(batch_X)\nexcept Exception as err:\n    print(err)\n    pred=[0.5 for _ in range(batch_size)]\npreds+=pred\nsub=pd.DataFrame()\nsub['filename']=os.listidr('../input/deepfake-detection-challenge/test_videos/')\nsub['label']=0.5\nfor pred,filename in zip(preds,filenames):\n    sub.loc[sub['filename']==filename,'label']=pred\nsub.to_csv('submission.csv',index=False)\n```\n\nYou can comment below if you are still experiencing those kind of errors after these fixes. Be sure to provide key parts of code.\n\nIf your error is timeout, do remember that the unseen validation set is 4000 videos, which is 10 times more. Your notebook should finish within 54 minutes(if we consider it as a linear relationship) when committing in order to finish predicting the unseen validation set. If not, be sure to turn on GPU(if that helps) and optimize your code to make it faster.",
    "730066": "You might also want to add error handling around the inference part of your code. For testing error handling on my pipeline I purposely corrupted one of the files to force it throw an error.",
    "730070": "ok thanks. But even if we added error handling, how will we be able to predict the rest. So the only thing we could do is to avoid  corrupted ones when preprocessing.",
    "730083": "If your using ffmpeg and subprocess calls a simple try/except will not catch corrupted file errors.\n\nsubprocess returns a value of 0 when the call is successful - when it fails other values are returned - so I used\n\na = subprocess.....\n\nif a ==0:\n    happyness\nelse:\n    sad - got here due to the error",
    "730484": "To clarify I mean add error handling when running inference on each batch, if there is an error running inference on the batch set results to 0.5 and continue processing the rest of the batches.",
    "730524": "Thanks. Added to the post.",
    "737869": "How is that even possible that the sample submission doesn't contain all files? Why this mess? Miss file names? What game is that? Maybe we don't want to play it..\n... Supposed to be data competition...",
    "782180": "Hi, I tried your error solution and the my submission finished in 1 hour, getting 0.68828, does this mean that I submitted all 0.5s?",
    "782215": "If the submission fails, the submission usually finishes in less than 30 minutes.\nCan you share the run time of your commit too."
  },
  "source": "meta"
}