{
  "id": 133225,
  "title": "|| processing",
  "url": "/competitions/deepfake-detection-challenge/discussion/133225",
  "author_name": "",
  "post_date": "2020-03-01T13:41:13.446290500Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Since the training is quite large, using the compute power at your disposal in an optimal fashion is very useful. Here I will share various code snippets to do some processing functions. I will improve these as I go. Feel free to comment and suggest better ways, thanks in advance. :) </p>\n\n<ul>\n<li>Unzipping the whole data</li>\n</ul>\n\n<p>```\nimport zipfile\nfrom pathlib import Path\nimport multiprocessing as mp</p>\n\n<h1>Inspired from this:</h1>\n\n<h1><a href=\"https://stackoverflow.com/questions/43313666/python-parallel-processing-to-unzip-files\">https://stackoverflow.com/questions/43313666/python-parallel-processing-to-unzip-files</a></h1>\n\n<h1>Change this to where your data lives.</h1>\n\n<p>BASE_PATH = Path(\"/hdd/deepfakes/train_videos\")</p>\n\n<p>def get_zip_files(base_path):\n    \"\"\" Get list of zip files. \"\"\"\n    return [f for f in base_path.glob(\"*.zip\")]</p>\n\n<p>def unzip(input_file):\n    \"\"\"Unzip a single file\"\"\"\n    with zipfile.ZipFile(input_file, 'r') as zip_ref:\n        zip_ref.extractall(BASE_PATH)</p>\n\n<p>def parallel_unzip(pool, zip_files):\n    \"\"\"Run an unzip process in //\"\"\"\n    pool.map(unzip, zip_files, chunksize=1)</p>\n\n<p>def delete(input_file):\n    \"\"\"Delete a single file\"\"\"\n    input_file.unlink()</p>\n\n<p>def parallel_delete(pool, zip_files):\n    \"\"\"Run a delete process in //\"\"\"\n    pool.map(delete, zip_files, chunksize=1)</p>\n\n<p>def pre_process_files(base_path, debug=True):\n    zip_files = get_zip_files(base_path)\n    if debug:\n        zip_files = [zip_files[0]]\n    cpus = min(mp.cpu_count(), len(zip_files))\n    pool = mp.Pool(cpus)\n    print(f\"Processing using {cpus} CPU workers\")\n    print(\"Start unzipping\")\n    parallel_unzip(pool, zip_files)\n    print(\"Done unzipping\")\n    print(\"Start deleting\")\n    parallel_delete(pool, zip_files)\n    print(\"Done deleting\")\n    pool.close()</p>\n\n<p>if <strong>name</strong> == \"<strong>main</strong>\":\n    #Make sure to test your script first before running it on all data. ;) \n    pre_process_files(BASE_PATH, debug=False)\n```</p>\n\n<p>In my case, 8 CPU workers running at the same time (see screenshot below): </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F1dd09cb0da7e844b23266c20a55e3afb%2FScreenshot%20from%202020-03-01%2014-44-28.png?generation=1583070302849413&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li>Extract faces + crop from a video using <a href=\"http://dlib.net/\">dlib</a> (not yet ||)</li>\n</ul>\n\n<p>```</p>\n\n<h1>Some of the code inspired from here:</h1>\n\n<h1>&nbsp;<a href=\"https://github.com/HongguLiu/Deepfake-Detection/blob/master/detect_from_video.py\">https://github.com/HongguLiu/Deepfake-Detection/blob/master/detect_from_video.py</a></h1>\n\n<p>import cv2\nimport dlib\nfrom pathlib import Path\nimport multiprocessing as mp\nfrom functools import partial</p>\n\n<h1>The face detector</h1>\n\n<p>face_detector = dlib.get_frontal_face_detector()</p>\n\n<h1>Change this folder to where you want to save the cropped faces.</h1>\n\n<p>FACE_EXTRACTED_FOLDER = Path(\"/hdd/deepfakes/face_extracted_from_train\")\nTRAIN_FOLDER = Path(\"/hdd/deepfakes/train_videos\")</p>\n\n<p>def get_boundingbox(face, width, height, scale=1.3, minsize=None):\n    x1 = face.left()\n    y1 = face.top()\n    x2 = face.right()\n    y2 = face.bottom()\n    size_bb = int(max(x2 - x1, y2 - y1) * scale)\n    if minsize:\n        if size_bb &lt; minsize:\n            size_bb = minsize\n    center_x, center_y = (x1 + x2) // 2, (y1 + y2) // 2</p>\n\n<pre><code># Check for out of bounds, x-y top left corner\nx1 = max(int(center_x - size_bb // 2), 0)\ny1 = max(int(center_y - size_bb // 2), 0)\n# Check for too big bb size for given x, y\nsize_bb = min(width - x1, size_bb)\nsize_bb = min(height - y1, size_bb)\n\nreturn x1, y1, size_bb\n</code></pre>\n\n<p>def detect_face_from_video(video_path, debug=True):\n    \"\"\"\n    Detect faces from a video and crop them. \n    \"\"\"</p>\n\n<pre><code># Pathify \nvideo_path = Path(video_path)\nprint(f\"Processing for {video_path.stem} in debug={debug} mode.\")\nreader = cv2.VideoCapture(video_path.as_posix())\nfile_name = video_path.stem\n# How many faces have been detected?\nfaces_counter = 0\nfolder = FACE_EXTRACTED_FOLDER / file_name\n# Create if it doesn't exist\nfolder.mkdir(parents=True, exist_ok=True)\n\n# TODO: Add a progress bar? Add a timer?\n\nwhile reader.isOpened():\n    _, image = reader.read()\n    try:\n        gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)\n    #&amp;nbsp;TODO: Investigate?\n    except Exception as e:\n        print(e)\n        break\n    _faces = face_detector(gray, 1)\n    if len(_faces):\n        height, width = image.shape[:2]\n        x, y, size = get_boundingbox(_faces[0], width, height)\n        cropped_face = image[y:y+size, x:x+size]\n        path = folder / f'{faces_counter}.png'\n        print(path)\n        cv2.imwrite(path.as_posix(), cropped_face)\n        if debug:\n            # Plot the cropped faces.\n            cv2.imshow(f\"cropped_image_{faces_counter}.png\", cropped_face)\n            cv2.waitKey(1000)\n            cv2.destroyAllWindows()\n        faces_counter += 1\nprint(f\"Have detected  {faces_counter} faces!\")\n</code></pre>\n\n<p>def parallel_detect_face_from_video(pool, files, debug):\n    \"\"\"Run a face detection from process in //\"\"\"\n    pool.map(partial(detect_face_from_video, debug=debug), files, chunksize=1)</p>\n\n<p>def main(debug=True):\n    video_paths = list(TRAIN_FOLDER.glob(\"*<em>/</em>.mp4\"))\n    if debug:\n        video_paths = [video_paths[0]]\n    num_cpus = min(mp.cpu_count(), len(video_paths))\n    print(\"Extracting faces using {} CPU workers.\")\n    with mp.Pool(num_cpus) as pool:\n        parallel_detect_face_from_video(pool, video_paths, debug=debug)</p>\n\n<p>```</p>\n\n<p>[EDIT 2-3-2020] Added a face extraction snippet (not yet ||)\n[EDIT 3-3-2020] Fixed a bug, added image plot section and image saving (still not yet ||). <br>\n[EDIT 3-3-2020] Now the face extraction bit is || as well. ;) \n[EDIT 8-3-2020] A new and better face extraction script can be found here =&gt; <a href=\"https://www.kaggle.com/yassinealouini/face-extraction/output?scriptVersionId=29839754\">https://www.kaggle.com/yassinealouini/face-extraction/output?scriptVersionId=29839754</a> </p>\n\n<p>Coming up next, predicting faces using a pretrained model, here is a sample: </p>\n\n<p>For a FAKE video (<strong>dfdc_train_part_0/owxbbpjpch.mp4</strong>), here are 300 predicted faces (1 if FAKE and 0 if TRUE): \n[0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0]</p>\n\n<p>Stay tuned. ;) </p>",
  "messages": [
    {
      "id": "760576",
      "postDate": "03/01/2020 13:41:13",
      "content": "<p>Since the training is quite large, using the compute power at your disposal in an optimal fashion is very useful. Here I will share various code snippets to do some processing functions. I will improve these as I go. Feel free to comment and suggest better ways, thanks in advance. :) </p>\n\n<ul>\n<li>Unzipping the whole data</li>\n</ul>\n\n<p>```\nimport zipfile\nfrom pathlib import Path\nimport multiprocessing as mp</p>\n\n<h1>Inspired from this:</h1>\n\n<h1><a href=\"https://stackoverflow.com/questions/43313666/python-parallel-processing-to-unzip-files\">https://stackoverflow.com/questions/43313666/python-parallel-processing-to-unzip-files</a></h1>\n\n<h1>Change this to where your data lives.</h1>\n\n<p>BASE_PATH = Path(\"/hdd/deepfakes/train_videos\")</p>\n\n<p>def get_zip_files(base_path):\n    \"\"\" Get list of zip files. \"\"\"\n    return [f for f in base_path.glob(\"*.zip\")]</p>\n\n<p>def unzip(input_file):\n    \"\"\"Unzip a single file\"\"\"\n    with zipfile.ZipFile(input_file, 'r') as zip_ref:\n        zip_ref.extractall(BASE_PATH)</p>\n\n<p>def parallel_unzip(pool, zip_files):\n    \"\"\"Run an unzip process in //\"\"\"\n    pool.map(unzip, zip_files, chunksize=1)</p>\n\n<p>def delete(input_file):\n    \"\"\"Delete a single file\"\"\"\n    input_file.unlink()</p>\n\n<p>def parallel_delete(pool, zip_files):\n    \"\"\"Run a delete process in //\"\"\"\n    pool.map(delete, zip_files, chunksize=1)</p>\n\n<p>def pre_process_files(base_path, debug=True):\n    zip_files = get_zip_files(base_path)\n    if debug:\n        zip_files = [zip_files[0]]\n    cpus = min(mp.cpu_count(), len(zip_files))\n    pool = mp.Pool(cpus)\n    print(f\"Processing using {cpus} CPU workers\")\n    print(\"Start unzipping\")\n    parallel_unzip(pool, zip_files)\n    print(\"Done unzipping\")\n    print(\"Start deleting\")\n    parallel_delete(pool, zip_files)\n    print(\"Done deleting\")\n    pool.close()</p>\n\n<p>if <strong>name</strong> == \"<strong>main</strong>\":\n    #Make sure to test your script first before running it on all data. ;) \n    pre_process_files(BASE_PATH, debug=False)\n```</p>\n\n<p>In my case, 8 CPU workers running at the same time (see screenshot below): </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F1dd09cb0da7e844b23266c20a55e3afb%2FScreenshot%20from%202020-03-01%2014-44-28.png?generation=1583070302849413&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li>Extract faces + crop from a video using <a href=\"http://dlib.net/\">dlib</a> (not yet ||)</li>\n</ul>\n\n<p>```</p>\n\n<h1>Some of the code inspired from here:</h1>\n\n<h1>&nbsp;<a href=\"https://github.com/HongguLiu/Deepfake-Detection/blob/master/detect_from_video.py\">https://github.com/HongguLiu/Deepfake-Detection/blob/master/detect_from_video.py</a></h1>\n\n<p>import cv2\nimport dlib\nfrom pathlib import Path\nimport multiprocessing as mp\nfrom functools import partial</p>\n\n<h1>The face detector</h1>\n\n<p>face_detector = dlib.get_frontal_face_detector()</p>\n\n<h1>Change this folder to where you want to save the cropped faces.</h1>\n\n<p>FACE_EXTRACTED_FOLDER = Path(\"/hdd/deepfakes/face_extracted_from_train\")\nTRAIN_FOLDER = Path(\"/hdd/deepfakes/train_videos\")</p>\n\n<p>def get_boundingbox(face, width, height, scale=1.3, minsize=None):\n    x1 = face.left()\n    y1 = face.top()\n    x2 = face.right()\n    y2 = face.bottom()\n    size_bb = int(max(x2 - x1, y2 - y1) * scale)\n    if minsize:\n        if size_bb &lt; minsize:\n            size_bb = minsize\n    center_x, center_y = (x1 + x2) // 2, (y1 + y2) // 2</p>\n\n<pre><code># Check for out of bounds, x-y top left corner\nx1 = max(int(center_x - size_bb // 2), 0)\ny1 = max(int(center_y - size_bb // 2), 0)\n# Check for too big bb size for given x, y\nsize_bb = min(width - x1, size_bb)\nsize_bb = min(height - y1, size_bb)\n\nreturn x1, y1, size_bb\n</code></pre>\n\n<p>def detect_face_from_video(video_path, debug=True):\n    \"\"\"\n    Detect faces from a video and crop them. \n    \"\"\"</p>\n\n<pre><code># Pathify \nvideo_path = Path(video_path)\nprint(f\"Processing for {video_path.stem} in debug={debug} mode.\")\nreader = cv2.VideoCapture(video_path.as_posix())\nfile_name = video_path.stem\n# How many faces have been detected?\nfaces_counter = 0\nfolder = FACE_EXTRACTED_FOLDER / file_name\n# Create if it doesn't exist\nfolder.mkdir(parents=True, exist_ok=True)\n\n# TODO: Add a progress bar? Add a timer?\n\nwhile reader.isOpened():\n    _, image = reader.read()\n    try:\n        gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)\n    #&amp;nbsp;TODO: Investigate?\n    except Exception as e:\n        print(e)\n        break\n    _faces = face_detector(gray, 1)\n    if len(_faces):\n        height, width = image.shape[:2]\n        x, y, size = get_boundingbox(_faces[0], width, height)\n        cropped_face = image[y:y+size, x:x+size]\n        path = folder / f'{faces_counter}.png'\n        print(path)\n        cv2.imwrite(path.as_posix(), cropped_face)\n        if debug:\n            # Plot the cropped faces.\n            cv2.imshow(f\"cropped_image_{faces_counter}.png\", cropped_face)\n            cv2.waitKey(1000)\n            cv2.destroyAllWindows()\n        faces_counter += 1\nprint(f\"Have detected  {faces_counter} faces!\")\n</code></pre>\n\n<p>def parallel_detect_face_from_video(pool, files, debug):\n    \"\"\"Run a face detection from process in //\"\"\"\n    pool.map(partial(detect_face_from_video, debug=debug), files, chunksize=1)</p>\n\n<p>def main(debug=True):\n    video_paths = list(TRAIN_FOLDER.glob(\"*<em>/</em>.mp4\"))\n    if debug:\n        video_paths = [video_paths[0]]\n    num_cpus = min(mp.cpu_count(), len(video_paths))\n    print(\"Extracting faces using {} CPU workers.\")\n    with mp.Pool(num_cpus) as pool:\n        parallel_detect_face_from_video(pool, video_paths, debug=debug)</p>\n\n<p>```</p>\n\n<p>[EDIT 2-3-2020] Added a face extraction snippet (not yet ||)\n[EDIT 3-3-2020] Fixed a bug, added image plot section and image saving (still not yet ||). <br>\n[EDIT 3-3-2020] Now the face extraction bit is || as well. ;) \n[EDIT 8-3-2020] A new and better face extraction script can be found here =&gt; <a href=\"https://www.kaggle.com/yassinealouini/face-extraction/output?scriptVersionId=29839754\">https://www.kaggle.com/yassinealouini/face-extraction/output?scriptVersionId=29839754</a> </p>\n\n<p>Coming up next, predicting faces using a pretrained model, here is a sample: </p>\n\n<p>For a FAKE video (<strong>dfdc_train_part_0/owxbbpjpch.mp4</strong>), here are 300 predicted faces (1 if FAKE and 0 if TRUE): \n[0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0]</p>\n\n<p>Stay tuned. ;) </p>",
      "rawMarkdown": "Since the training is quite large, using the compute power at your disposal in an optimal fashion is very useful. Here I will share various code snippets to do some processing functions. I will improve these as I go. Feel free to comment and suggest better ways, thanks in advance. :) \n\n\n\n- Unzipping the whole data\n\n```\nimport zipfile\nfrom pathlib import Path\nimport multiprocessing as mp\n\n# Inspired from this: \n# https://stackoverflow.com/questions/43313666/python-parallel-processing-to-unzip-files\n\n# Change this to where your data lives. \nBASE_PATH = Path(\"/hdd/deepfakes/train_videos\")\n\n\ndef get_zip_files(base_path):\n    \"\"\" Get list of zip files. \"\"\"\n    return [f for f in base_path.glob(\"*.zip\")]\n\ndef unzip(input_file):\n    \"\"\"Unzip a single file\"\"\"\n    with zipfile.ZipFile(input_file, 'r') as zip_ref:\n        zip_ref.extractall(BASE_PATH)\n\ndef parallel_unzip(pool, zip_files):\n    \"\"\"Run an unzip process in //\"\"\"\n    pool.map(unzip, zip_files, chunksize=1)\n\ndef delete(input_file):\n    \"\"\"Delete a single file\"\"\"\n    input_file.unlink()\n\ndef parallel_delete(pool, zip_files):\n    \"\"\"Run a delete process in //\"\"\"\n    pool.map(delete, zip_files, chunksize=1)\n    \n\ndef pre_process_files(base_path, debug=True):\n    zip_files = get_zip_files(base_path)\n    if debug:\n        zip_files = [zip_files[0]]\n    cpus = min(mp.cpu_count(), len(zip_files))\n    pool = mp.Pool(cpus)\n    print(f\"Processing using {cpus} CPU workers\")\n    print(\"Start unzipping\")\n    parallel_unzip(pool, zip_files)\n    print(\"Done unzipping\")\n    print(\"Start deleting\")\n    parallel_delete(pool, zip_files)\n    print(\"Done deleting\")\n    pool.close()\n\n\nif __name__ == \"__main__\":\n    #Make sure to test your script first before running it on all data. ;) \n    pre_process_files(BASE_PATH, debug=False)\n```\n\n\nIn my case, 8 CPU workers running at the same time (see screenshot below): \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F1dd09cb0da7e844b23266c20a55e3afb%2FScreenshot%20from%202020-03-01%2014-44-28.png?generation=1583070302849413&amp;alt=media)\n\n\n- Extract faces + crop from a video using [dlib](http://dlib.net/) (not yet ||)\n\n```\n\n# Some of the code inspired from here: \n#&nbsp;https://github.com/HongguLiu/Deepfake-Detection/blob/master/detect_from_video.py\n\nimport cv2\nimport dlib\nfrom pathlib import Path\nimport multiprocessing as mp\nfrom functools import partial\n\n# The face detector\nface_detector = dlib.get_frontal_face_detector()\n\n# Change this folder to where you want to save the cropped faces. \nFACE_EXTRACTED_FOLDER = Path(\"/hdd/deepfakes/face_extracted_from_train\")\nTRAIN_FOLDER = Path(\"/hdd/deepfakes/train_videos\")\n\n\n\n\ndef get_boundingbox(face, width, height, scale=1.3, minsize=None):\n    x1 = face.left()\n    y1 = face.top()\n    x2 = face.right()\n    y2 = face.bottom()\n    size_bb = int(max(x2 - x1, y2 - y1) * scale)\n    if minsize:\n        if size_bb &lt; minsize:\n            size_bb = minsize\n    center_x, center_y = (x1 + x2) // 2, (y1 + y2) // 2\n\n    # Check for out of bounds, x-y top left corner\n    x1 = max(int(center_x - size_bb // 2), 0)\n    y1 = max(int(center_y - size_bb // 2), 0)\n    # Check for too big bb size for given x, y\n    size_bb = min(width - x1, size_bb)\n    size_bb = min(height - y1, size_bb)\n\n    return x1, y1, size_bb\n\n\ndef detect_face_from_video(video_path, debug=True):\n    \"\"\"\n    Detect faces from a video and crop them. \n    \"\"\"\n\n\n    # Pathify \n    video_path = Path(video_path)\n    print(f\"Processing for {video_path.stem} in debug={debug} mode.\")\n    reader = cv2.VideoCapture(video_path.as_posix())\n    file_name = video_path.stem\n    # How many faces have been detected?\n    faces_counter = 0\n    folder = FACE_EXTRACTED_FOLDER / file_name\n    # Create if it doesn't exist\n    folder.mkdir(parents=True, exist_ok=True)\n\n    # TODO: Add a progress bar? Add a timer?\n\n    while reader.isOpened():\n        _, image = reader.read()\n        try:\n            gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)\n        #&nbsp;TODO: Investigate?\n        except Exception as e:\n            print(e)\n            break\n        _faces = face_detector(gray, 1)\n        if len(_faces):\n            height, width = image.shape[:2]\n            x, y, size = get_boundingbox(_faces[0], width, height)\n            cropped_face = image[y:y+size, x:x+size]\n            path = folder / f'{faces_counter}.png'\n            print(path)\n            cv2.imwrite(path.as_posix(), cropped_face)\n            if debug:\n                # Plot the cropped faces.\n                cv2.imshow(f\"cropped_image_{faces_counter}.png\", cropped_face)\n                cv2.waitKey(1000)\n                cv2.destroyAllWindows()\n            faces_counter += 1\n    print(f\"Have detected  {faces_counter} faces!\")\n\n\ndef parallel_detect_face_from_video(pool, files, debug):\n    \"\"\"Run a face detection from process in //\"\"\"\n    pool.map(partial(detect_face_from_video, debug=debug), files, chunksize=1)\n\n\ndef main(debug=True):\n    video_paths = list(TRAIN_FOLDER.glob(\"**/*.mp4\"))\n    if debug:\n        video_paths = [video_paths[0]]\n    num_cpus = min(mp.cpu_count(), len(video_paths))\n    print(\"Extracting faces using {} CPU workers.\")\n    with mp.Pool(num_cpus) as pool:\n        parallel_detect_face_from_video(pool, video_paths, debug=debug)\n\n\n```\n\n\n[EDIT 2-3-2020] Added a face extraction snippet (not yet ||)\n[EDIT 3-3-2020] Fixed a bug, added image plot section and image saving (still not yet ||).  \n[EDIT 3-3-2020] Now the face extraction bit is || as well. ;) \n[EDIT 8-3-2020] A new and better face extraction script can be found here =&gt; https://www.kaggle.com/yassinealouini/face-extraction/output?scriptVersionId=29839754 \n\nComing up next, predicting faces using a pretrained model, here is a sample: \n\nFor a FAKE video (**dfdc_train_part_0/owxbbpjpch.mp4**), here are 300 predicted faces (1 if FAKE and 0 if TRUE): \n[0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0]\n\nStay tuned. ;)",
      "votes": null
    },
    {
      "id": "760583",
      "postDate": "03/01/2020 13:45:26",
      "content": "<p>Awesome!</p>\n\n<p>One day, your tips shall lead to your fall tho :-(</p>",
      "rawMarkdown": "Awesome!\n\nOne day, your tips shall lead to your fall tho :-(",
      "votes": null
    },
    {
      "id": "760586",
      "postDate": "03/01/2020 13:47:45",
      "content": "<p>Thanks! No worries, I am here to learn and share, competing isn't my main objective. ;) </p>",
      "rawMarkdown": "Thanks! No worries, I am here to learn and share, competing isn't my main objective. ;)",
      "votes": null
    },
    {
      "id": "760587",
      "postDate": "03/01/2020 13:49:02",
      "content": "<p>Although, I have a feeling whoever sees this will want to merge this with the top public kernel.</p>",
      "rawMarkdown": "Although, I have a feeling whoever sees this will want to merge this with the top public kernel.",
      "votes": null
    },
    {
      "id": "760685",
      "postDate": "03/01/2020 16:05:09",
      "content": "<p>Some tricks for dealing with a lot of files/folders:</p>\n\n<ul>\n<li>Prefer <code>pathlib</code> over <code>os</code>, it has an easier/nicer API (this is a personal taste).</li>\n<li>Test your script on a folder (or even a single file) before running on everything. </li>\n<li><code>multiprocessing</code> is often key to || processing. </li>\n</ul>",
      "rawMarkdown": "Some tricks for dealing with a lot of files/folders:\n\n- Prefer `pathlib` over `os`, it has an easier/nicer API (this is a personal taste).\n- Test your script on a folder (or even a single file) before running on everything. \n-  `multiprocessing` is often key to || processing.",
      "votes": null
    },
    {
      "id": "760799",
      "postDate": "03/01/2020 18:26:28",
      "content": "<p>While I am writing the face extraction || processing part, here is one of the labeled frames from running this script: </p>\n\n<p><code>python detect_from_video.py --video_path ../train_videos/dfdc_train_part_0/wkczijuamz.mp4 --model_path ./pretrained_model/df_c0_best.pkl -o /hdd/deepfakes/face_extracted_from_train</code></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F8f085451a550ba0808795c79a7321403%2FScreenshot%20from%202020-03-01%2019-21-48.png?generation=1583087081558892&amp;alt=media\" alt=\"\"></p>\n\n<p>For instance, this is a FAKE video (as demonstrated from the metadata file):</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F36a9317e5c75b8ad96e10168667ee932%2FScreenshot%20from%202020-03-01%2019-24-55.png?generation=1583087132376566&amp;alt=media\" alt=\"\"></p>\n\n<p>The script comes from this github repo: <a href=\"https://github.com/HongguLiu/Deepfake-Detection/blob/master/detect_from_video.py\">https://github.com/HongguLiu/Deepfake-Detection/blob/master/detect_from_video.py</a>.</p>",
      "rawMarkdown": "While I am writing the face extraction || processing part, here is one of the labeled frames from running this script: \n\n`python detect_from_video.py --video_path ../train_videos/dfdc_train_part_0/wkczijuamz.mp4 --model_path ./pretrained_model/df_c0_best.pkl -o /hdd/deepfakes/face_extracted_from_train`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F8f085451a550ba0808795c79a7321403%2FScreenshot%20from%202020-03-01%2019-21-48.png?generation=1583087081558892&amp;alt=media)\n\nFor instance, this is a FAKE video (as demonstrated from the metadata file):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F36a9317e5c75b8ad96e10168667ee932%2FScreenshot%20from%202020-03-01%2019-24-55.png?generation=1583087132376566&amp;alt=media)\n\nThe script comes from this github repo: https://github.com/HongguLiu/Deepfake-Detection/blob/master/detect_from_video.py.",
      "votes": null
    },
    {
      "id": "762860",
      "postDate": "03/03/2020 21:58:46",
      "content": "<p>Any opencv experts here that can help me with this message error (could use Google but better to ask here I guess ;)):</p>\n\n<p><code>OpenCV(4.2.0) /io/opencv/modules/imgproc/src/color.cpp:182: error: (-215:Assertion failed) !_src.empty() in function 'cvtColor'</code></p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Any opencv experts here that can help me with this message error (could use Google but better to ask here I guess ;)):\n\n`OpenCV(4.2.0) /io/opencv/modules/imgproc/src/color.cpp:182: error: (-215:Assertion failed) !_src.empty() in function 'cvtColor'`\n\nThanks!",
      "votes": null
    },
    {
      "id": "762861",
      "postDate": "03/03/2020 21:59:26",
      "content": "<p>That what be interesting for sure!</p>",
      "rawMarkdown": "That what be interesting for sure!",
      "votes": null
    },
    {
      "id": "762884",
      "postDate": "03/03/2020 22:40:31",
      "content": "<p>|| face extraction in action, time to sleep I guess. ;)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2Fd512807be71bb9774b0531fdddb839b4%2FScreenshot%20from%202020-03-03%2023-39-41.png?generation=1583275206166615&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2Fad5caeda84dc5c06f9948dbd3cc89960%2FScreenshot%20from%202020-03-03%2023-39-34.png?generation=1583275208909226&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "|| face extraction in action, time to sleep I guess. ;)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2Fd512807be71bb9774b0531fdddb839b4%2FScreenshot%20from%202020-03-03%2023-39-41.png?generation=1583275206166615&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2Fad5caeda84dc5c06f9948dbd3cc89960%2FScreenshot%20from%202020-03-03%2023-39-34.png?generation=1583275208909226&amp;alt=media)",
      "votes": null
    },
    {
      "id": "763037",
      "postDate": "03/04/2020 03:56:29",
      "content": "<p>It means the image is empty.</p>",
      "rawMarkdown": "It means the image is empty.",
      "votes": null
    },
    {
      "id": "766646",
      "postDate": "03/08/2020 14:09:26",
      "content": "<p>Here is a link to a script that performs face extraction (faster and smarter thanks to the many suggestions ;)) =&gt; <a href=\"https://www.kaggle.com/yassinealouini/face-extraction/output?scriptVersionId=29839754\">https://www.kaggle.com/yassinealouini/face-extraction/output?scriptVersionId=29839754</a>. Three levels of optimization (at least): </p>\n\n<ul>\n<li>Extract bounding boxes only for REAL faces and save these =&gt; thus running for 19154 videos instead of the whole 119146 files =&gt; an 84% gain!</li>\n<li>Skip some of the frames when extracting faces =&gt; the more frames one skipps, the faster it runs. </li>\n<li>|| execution for I/O operation using multiprocessing</li>\n</ul>\n\n<p>There is only one step missing: cropping FAKE videos using the extracted faces bounding boxes. I won't share this step so that I don't reveal too much (this step is easy to figure). </p>\n\n<p>Also, other features can be added to the bounding boxes Parquet file to deal with \"funny\" artifacts (as mentioned <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134352\">here</a>) in FAKE videos. </p>\n\n<p>Let me know if this is of any help and if you have other tips to suggest. </p>",
      "rawMarkdown": "Here is a link to a script that performs face extraction (faster and smarter thanks to the many suggestions ;)) =&gt; https://www.kaggle.com/yassinealouini/face-extraction/output?scriptVersionId=29839754. Three levels of optimization (at least): \n\n- Extract bounding boxes only for REAL faces and save these =&gt; thus running for 19154 videos instead of the whole 119146 files =&gt; an 84% gain!\n- Skip some of the frames when extracting faces =&gt; the more frames one skipps, the faster it runs. \n- || execution for I/O operation using multiprocessing\n\nThere is only one step missing: cropping FAKE videos using the extracted faces bounding boxes. I won't share this step so that I don't reveal too much (this step is easy to figure). \n\nAlso, other features can be added to the bounding boxes Parquet file to deal with \"funny\" artifacts (as mentioned [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134352)) in FAKE videos. \n\nLet me know if this is of any help and if you have other tips to suggest.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 760583,
      "author_name": "nxrprime",
      "author_url": "",
      "post_date": "03/01/2020 13:45:26",
      "content": "<p>Awesome!</p>\n\n<p>One day, your tips shall lead to your fall tho :-(</p>",
      "votes": null,
      "replies": [
        {
          "id": 760586,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "03/01/2020 13:47:45",
          "content": "<p>Thanks! No worries, I am here to learn and share, competing isn't my main objective. ;) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760587,
          "author_name": "nxrprime",
          "author_url": "",
          "post_date": "03/01/2020 13:49:02",
          "content": "<p>Although, I have a feeling whoever sees this will want to merge this with the top public kernel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 762861,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "03/03/2020 21:59:26",
          "content": "<p>That what be interesting for sure!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 760685,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "03/01/2020 16:05:09",
      "content": "<p>Some tricks for dealing with a lot of files/folders:</p>\n\n<ul>\n<li>Prefer <code>pathlib</code> over <code>os</code>, it has an easier/nicer API (this is a personal taste).</li>\n<li>Test your script on a folder (or even a single file) before running on everything. </li>\n<li><code>multiprocessing</code> is often key to || processing. </li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 760799,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "03/01/2020 18:26:28",
      "content": "<p>While I am writing the face extraction || processing part, here is one of the labeled frames from running this script: </p>\n\n<p><code>python detect_from_video.py --video_path ../train_videos/dfdc_train_part_0/wkczijuamz.mp4 --model_path ./pretrained_model/df_c0_best.pkl -o /hdd/deepfakes/face_extracted_from_train</code></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F8f085451a550ba0808795c79a7321403%2FScreenshot%20from%202020-03-01%2019-21-48.png?generation=1583087081558892&amp;alt=media\" alt=\"\"></p>\n\n<p>For instance, this is a FAKE video (as demonstrated from the metadata file):</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F36a9317e5c75b8ad96e10168667ee932%2FScreenshot%20from%202020-03-01%2019-24-55.png?generation=1583087132376566&amp;alt=media\" alt=\"\"></p>\n\n<p>The script comes from this github repo: <a href=\"https://github.com/HongguLiu/Deepfake-Detection/blob/master/detect_from_video.py\">https://github.com/HongguLiu/Deepfake-Detection/blob/master/detect_from_video.py</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 762860,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "03/03/2020 21:58:46",
      "content": "<p>Any opencv experts here that can help me with this message error (could use Google but better to ask here I guess ;)):</p>\n\n<p><code>OpenCV(4.2.0) /io/opencv/modules/imgproc/src/color.cpp:182: error: (-215:Assertion failed) !_src.empty() in function 'cvtColor'</code></p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 763037,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "03/04/2020 03:56:29",
          "content": "<p>It means the image is empty.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 762884,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "03/03/2020 22:40:31",
      "content": "<p>|| face extraction in action, time to sleep I guess. ;)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2Fd512807be71bb9774b0531fdddb839b4%2FScreenshot%20from%202020-03-03%2023-39-41.png?generation=1583275206166615&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2Fad5caeda84dc5c06f9948dbd3cc89960%2FScreenshot%20from%202020-03-03%2023-39-34.png?generation=1583275208909226&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 766646,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "03/08/2020 14:09:26",
      "content": "<p>Here is a link to a script that performs face extraction (faster and smarter thanks to the many suggestions ;)) =&gt; <a href=\"https://www.kaggle.com/yassinealouini/face-extraction/output?scriptVersionId=29839754\">https://www.kaggle.com/yassinealouini/face-extraction/output?scriptVersionId=29839754</a>. Three levels of optimization (at least): </p>\n\n<ul>\n<li>Extract bounding boxes only for REAL faces and save these =&gt; thus running for 19154 videos instead of the whole 119146 files =&gt; an 84% gain!</li>\n<li>Skip some of the frames when extracting faces =&gt; the more frames one skipps, the faster it runs. </li>\n<li>|| execution for I/O operation using multiprocessing</li>\n</ul>\n\n<p>There is only one step missing: cropping FAKE videos using the extracted faces bounding boxes. I won't share this step so that I don't reveal too much (this step is easy to figure). </p>\n\n<p>Also, other features can be added to the bounding boxes Parquet file to deal with \"funny\" artifacts (as mentioned <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134352\">here</a>) in FAKE videos. </p>\n\n<p>Let me know if this is of any help and if you have other tips to suggest. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "760576": "Since the training is quite large, using the compute power at your disposal in an optimal fashion is very useful. Here I will share various code snippets to do some processing functions. I will improve these as I go. Feel free to comment and suggest better ways, thanks in advance. :) \n\n\n\n- Unzipping the whole data\n\n```\nimport zipfile\nfrom pathlib import Path\nimport multiprocessing as mp\n\n# Inspired from this: \n# https://stackoverflow.com/questions/43313666/python-parallel-processing-to-unzip-files\n\n# Change this to where your data lives. \nBASE_PATH = Path(\"/hdd/deepfakes/train_videos\")\n\n\ndef get_zip_files(base_path):\n    \"\"\" Get list of zip files. \"\"\"\n    return [f for f in base_path.glob(\"*.zip\")]\n\ndef unzip(input_file):\n    \"\"\"Unzip a single file\"\"\"\n    with zipfile.ZipFile(input_file, 'r') as zip_ref:\n        zip_ref.extractall(BASE_PATH)\n\ndef parallel_unzip(pool, zip_files):\n    \"\"\"Run an unzip process in //\"\"\"\n    pool.map(unzip, zip_files, chunksize=1)\n\ndef delete(input_file):\n    \"\"\"Delete a single file\"\"\"\n    input_file.unlink()\n\ndef parallel_delete(pool, zip_files):\n    \"\"\"Run a delete process in //\"\"\"\n    pool.map(delete, zip_files, chunksize=1)\n    \n\ndef pre_process_files(base_path, debug=True):\n    zip_files = get_zip_files(base_path)\n    if debug:\n        zip_files = [zip_files[0]]\n    cpus = min(mp.cpu_count(), len(zip_files))\n    pool = mp.Pool(cpus)\n    print(f\"Processing using {cpus} CPU workers\")\n    print(\"Start unzipping\")\n    parallel_unzip(pool, zip_files)\n    print(\"Done unzipping\")\n    print(\"Start deleting\")\n    parallel_delete(pool, zip_files)\n    print(\"Done deleting\")\n    pool.close()\n\n\nif __name__ == \"__main__\":\n    #Make sure to test your script first before running it on all data. ;) \n    pre_process_files(BASE_PATH, debug=False)\n```\n\n\nIn my case, 8 CPU workers running at the same time (see screenshot below): \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F1dd09cb0da7e844b23266c20a55e3afb%2FScreenshot%20from%202020-03-01%2014-44-28.png?generation=1583070302849413&amp;alt=media)\n\n\n- Extract faces + crop from a video using [dlib](http://dlib.net/) (not yet ||)\n\n```\n\n# Some of the code inspired from here: \n#&nbsp;https://github.com/HongguLiu/Deepfake-Detection/blob/master/detect_from_video.py\n\nimport cv2\nimport dlib\nfrom pathlib import Path\nimport multiprocessing as mp\nfrom functools import partial\n\n# The face detector\nface_detector = dlib.get_frontal_face_detector()\n\n# Change this folder to where you want to save the cropped faces. \nFACE_EXTRACTED_FOLDER = Path(\"/hdd/deepfakes/face_extracted_from_train\")\nTRAIN_FOLDER = Path(\"/hdd/deepfakes/train_videos\")\n\n\n\n\ndef get_boundingbox(face, width, height, scale=1.3, minsize=None):\n    x1 = face.left()\n    y1 = face.top()\n    x2 = face.right()\n    y2 = face.bottom()\n    size_bb = int(max(x2 - x1, y2 - y1) * scale)\n    if minsize:\n        if size_bb &lt; minsize:\n            size_bb = minsize\n    center_x, center_y = (x1 + x2) // 2, (y1 + y2) // 2\n\n    # Check for out of bounds, x-y top left corner\n    x1 = max(int(center_x - size_bb // 2), 0)\n    y1 = max(int(center_y - size_bb // 2), 0)\n    # Check for too big bb size for given x, y\n    size_bb = min(width - x1, size_bb)\n    size_bb = min(height - y1, size_bb)\n\n    return x1, y1, size_bb\n\n\ndef detect_face_from_video(video_path, debug=True):\n    \"\"\"\n    Detect faces from a video and crop them. \n    \"\"\"\n\n\n    # Pathify \n    video_path = Path(video_path)\n    print(f\"Processing for {video_path.stem} in debug={debug} mode.\")\n    reader = cv2.VideoCapture(video_path.as_posix())\n    file_name = video_path.stem\n    # How many faces have been detected?\n    faces_counter = 0\n    folder = FACE_EXTRACTED_FOLDER / file_name\n    # Create if it doesn't exist\n    folder.mkdir(parents=True, exist_ok=True)\n\n    # TODO: Add a progress bar? Add a timer?\n\n    while reader.isOpened():\n        _, image = reader.read()\n        try:\n            gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)\n        #&nbsp;TODO: Investigate?\n        except Exception as e:\n            print(e)\n            break\n        _faces = face_detector(gray, 1)\n        if len(_faces):\n            height, width = image.shape[:2]\n            x, y, size = get_boundingbox(_faces[0], width, height)\n            cropped_face = image[y:y+size, x:x+size]\n            path = folder / f'{faces_counter}.png'\n            print(path)\n            cv2.imwrite(path.as_posix(), cropped_face)\n            if debug:\n                # Plot the cropped faces.\n                cv2.imshow(f\"cropped_image_{faces_counter}.png\", cropped_face)\n                cv2.waitKey(1000)\n                cv2.destroyAllWindows()\n            faces_counter += 1\n    print(f\"Have detected  {faces_counter} faces!\")\n\n\ndef parallel_detect_face_from_video(pool, files, debug):\n    \"\"\"Run a face detection from process in //\"\"\"\n    pool.map(partial(detect_face_from_video, debug=debug), files, chunksize=1)\n\n\ndef main(debug=True):\n    video_paths = list(TRAIN_FOLDER.glob(\"**/*.mp4\"))\n    if debug:\n        video_paths = [video_paths[0]]\n    num_cpus = min(mp.cpu_count(), len(video_paths))\n    print(\"Extracting faces using {} CPU workers.\")\n    with mp.Pool(num_cpus) as pool:\n        parallel_detect_face_from_video(pool, video_paths, debug=debug)\n\n\n```\n\n\n[EDIT 2-3-2020] Added a face extraction snippet (not yet ||)\n[EDIT 3-3-2020] Fixed a bug, added image plot section and image saving (still not yet ||).  \n[EDIT 3-3-2020] Now the face extraction bit is || as well. ;) \n[EDIT 8-3-2020] A new and better face extraction script can be found here =&gt; https://www.kaggle.com/yassinealouini/face-extraction/output?scriptVersionId=29839754 \n\nComing up next, predicting faces using a pretrained model, here is a sample: \n\nFor a FAKE video (**dfdc_train_part_0/owxbbpjpch.mp4**), here are 300 predicted faces (1 if FAKE and 0 if TRUE): \n[0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0]\n\nStay tuned. ;)",
    "760583": "Awesome!\n\nOne day, your tips shall lead to your fall tho :-(",
    "760586": "Thanks! No worries, I am here to learn and share, competing isn't my main objective. ;)",
    "760587": "Although, I have a feeling whoever sees this will want to merge this with the top public kernel.",
    "760685": "Some tricks for dealing with a lot of files/folders:\n\n- Prefer `pathlib` over `os`, it has an easier/nicer API (this is a personal taste).\n- Test your script on a folder (or even a single file) before running on everything. \n-  `multiprocessing` is often key to || processing.",
    "760799": "While I am writing the face extraction || processing part, here is one of the labeled frames from running this script: \n\n`python detect_from_video.py --video_path ../train_videos/dfdc_train_part_0/wkczijuamz.mp4 --model_path ./pretrained_model/df_c0_best.pkl -o /hdd/deepfakes/face_extracted_from_train`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F8f085451a550ba0808795c79a7321403%2FScreenshot%20from%202020-03-01%2019-21-48.png?generation=1583087081558892&amp;alt=media)\n\nFor instance, this is a FAKE video (as demonstrated from the metadata file):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F36a9317e5c75b8ad96e10168667ee932%2FScreenshot%20from%202020-03-01%2019-24-55.png?generation=1583087132376566&amp;alt=media)\n\nThe script comes from this github repo: https://github.com/HongguLiu/Deepfake-Detection/blob/master/detect_from_video.py.",
    "762860": "Any opencv experts here that can help me with this message error (could use Google but better to ask here I guess ;)):\n\n`OpenCV(4.2.0) /io/opencv/modules/imgproc/src/color.cpp:182: error: (-215:Assertion failed) !_src.empty() in function 'cvtColor'`\n\nThanks!",
    "762861": "That what be interesting for sure!",
    "762884": "|| face extraction in action, time to sleep I guess. ;)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2Fd512807be71bb9774b0531fdddb839b4%2FScreenshot%20from%202020-03-03%2023-39-41.png?generation=1583275206166615&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2Fad5caeda84dc5c06f9948dbd3cc89960%2FScreenshot%20from%202020-03-03%2023-39-34.png?generation=1583275208909226&amp;alt=media)",
    "763037": "It means the image is empty.",
    "766646": "Here is a link to a script that performs face extraction (faster and smarter thanks to the many suggestions ;)) =&gt; https://www.kaggle.com/yassinealouini/face-extraction/output?scriptVersionId=29839754. Three levels of optimization (at least): \n\n- Extract bounding boxes only for REAL faces and save these =&gt; thus running for 19154 videos instead of the whole 119146 files =&gt; an 84% gain!\n- Skip some of the frames when extracting faces =&gt; the more frames one skipps, the faster it runs. \n- || execution for I/O operation using multiprocessing\n\nThere is only one step missing: cropping FAKE videos using the extracted faces bounding boxes. I won't share this step so that I don't reveal too much (this step is easy to figure). \n\nAlso, other features can be added to the bounding boxes Parquet file to deal with \"funny\" artifacts (as mentioned [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134352)) in FAKE videos. \n\nLet me know if this is of any help and if you have other tips to suggest."
  },
  "source": "meta"
}