{"cells":[{"metadata":{},"cell_type":"markdown","source":"# What's wrong with this submission\nI forked [this kernel](https://www.kaggle.com/zaharch/public-test-errors) from @zaharch. I changed the last part to confirm if all of  corrupted videos are fake. But always got the \"Submission Scoring Error\".\n**Anyone knowns the reason. Thanks.**"},{"metadata":{},"cell_type":"markdown","source":"# 27 public test videos fail on conversion to PIL image"},{"metadata":{},"cell_type":"markdown","source":"I had several spare submissions so I tried to hunt down the public test errors that were plaguing the village. The conclusion so far is that there are 27 corrupted videos in the public test, and the failures for me are specifically at `Image.fromarray` call. Some more details below."},{"metadata":{"trusted":true},"cell_type":"code","source":"import os\nimport glob\nimport cv2\nfrom PIL import Image\nimport pandas as pd\nimport numpy as np","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"filenames = glob.glob('/kaggle/input/deepfake-detection-challenge/test_videos/*.mp4')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission = pd.read_csv(\"/kaggle/input/deepfake-detection-challenge/sample_submission.csv\", index_col=0)\nsubmission['label'] = 0.5","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(len(filenames))\nfor filename in filenames:\n    try:\n        name = os.path.basename(filename)\n        v_cap = cv2.VideoCapture(filename)\n        v_len = int(v_cap.get(cv2.CAP_PROP_FRAME_COUNT))\n        for j in range(v_len):\n            success = v_cap.grab()\n            if j == (v_len-1):\n                success, vframe = v_cap.retrieve()\n                vframe = Image.fromarray(vframe) # this line fails for 27 videos\n        v_cap.release()\n        submission.loc[name, 'label'] = 0.5\n    except:\n        submission.loc[name, 'label'] = 1\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission.to_csv('submission.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In the code above I am submitting the same values for all videos, and we know that there are exactly 2000 fakes and 2000 real videos in the public test, so the expected score we can calculate with this function:"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}