{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This notebook aims to demonstrate the different ways to use the MTCNN face detection module of `facenet-pytorch`. Originally reported in [Joint Face Detection and Alignment using Multi-task Cascaded Convolutional Networks](https://arxiv.org/abs/1604.02878), the MTCNN network is able to simultaneously propose bounding boxes, five-point facial landmarks, and detection probabilities. Taken from the original paper:\n\n> Face detection and alignment in unconstrained environments are challenging due to various poses, illuminations and occlusions. Recent studies show that deep learning approaches can achieve impressive performance on these two tasks. In this paper, we propose a deep cascaded multi-task framework which exploits the inherent correlation between them to boost up their performance. In particular, our framework adopts a cascaded structure with three stages of carefully designed deep convolutional networks that predict face and landmark location in a coarse-to-fine manner. In addition, in the learning process, we propose a new online hard sample mining strategy that can improve the performance automatically without manual sample selection. Our method achieves superior accuracy over the state-of-the-art techniques on the challenging FDDB and WIDER FACE benchmark for face detection, and AFLW benchmark for face alignment, while keeps real time performance.\n\n`facenet-pytorch` includes an efficient, cuda-ready implementation of MTCNN that will be demonstrated in this notebook. The following topics will be covered:\n\n1. <a href='#1'>Documentation</a>\n1. <a href='#2'>Basic usage</a>\n1. <a href='#3'>Preventing image normalization</a>\n1. <a href='#4'>Margin adjustment</a>\n1. <a href='#5'>Multiple faces in a single image</a>\n1. <a href='#6'>Batched detection</a>\n1. <a href='#7'>Bounding boxes and facial landmarks</a>\n1. <a href='#8'>Saving face datasets</a>\n\nOther resources:\n\n1. The facenet-pytorch [github repo](https://github.com/timesler/facenet-pytorch)\n1. [Notebook demonstrating combined use of face detection and recognition](https://www.kaggle.com/timesler/facial-recognition-model-in-pytorch)\n1. [The FastMTCNN algorithm](https://www.kaggle.com/timesler/fast-mtcnn-detector-45-fps-at-full-resolution) ","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install ../input/facenet-pytorch-vggface2/facenet_pytorch-2.2.9-py3-none-any.whl","metadata":{"execution":{"iopub.status.busy":"2022-03-30T07:36:59.738267Z","iopub.execute_input":"2022-03-30T07:36:59.738572Z","iopub.status.idle":"2022-03-30T07:37:26.680872Z","shell.execute_reply.started":"2022-03-30T07:36:59.738506Z","shell.execute_reply":"2022-03-30T07:37:26.680030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from facenet_pytorch import MTCNN\nimport cv2\nfrom PIL import Image\nimport numpy as np\nfrom matplotlib import pyplot as plt\nfrom tqdm.notebook import tqdm","metadata":{"execution":{"iopub.status.busy":"2022-03-30T07:37:34.212511Z","iopub.execute_input":"2022-03-30T07:37:34.212846Z","iopub.status.idle":"2022-03-30T07:37:35.690428Z","shell.execute_reply.started":"2022-03-30T07:37:34.212785Z","shell.execute_reply":"2022-03-30T07:37:35.689712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a id='1'>Documentation</a>\n\nDetailed usage information is contained in the MTCNN docstring:\n\n```\nhelp(MTCNN)\n```","metadata":{}},{"cell_type":"code","source":"help(MTCNN)","metadata":{"execution":{"iopub.status.busy":"2022-03-30T03:00:21.438866Z","iopub.execute_input":"2022-03-30T03:00:21.439168Z","iopub.status.idle":"2022-03-30T03:00:21.452574Z","shell.execute_reply.started":"2022-03-30T03:00:21.439116Z","shell.execute_reply":"2022-03-30T03:00:21.451826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a id='2'>Basic usage</a>\n\nUnlike other implementations, calling a `facenet-pytorch` MTCNN object directly with an image (i.e., using the forward method for those familiar with pytorch) will return torch tensors containing the detected face(s), rather than just the bounding boxes. This is to enable using the module easily as the first stage of a facial recognition pipeline, in which the faces are passed directly to an additional network or algorithm.\n\nIn order to return the detected boxes instead (and optionally, the facial landmarks), see the `MTCNN.detect()` method. Its use will be described below also.\n\nTo create an MTCNN detector that runs on the GPU, instantiate the model with `device='cuda:0'` or equivalent.\n\nFor this competition, it will be best to set `select_largest=False` to ensure detected faces are ordered according to detection probability rather than size.","metadata":{}},{"cell_type":"code","source":"# Create face detector\nmtcnn = MTCNN(select_largest=False, device='cuda')# face with the highest detection  probability is return\n\n# Load a single image and display\nv_cap = cv2.VideoCapture('/kaggle/input/deepfake-detection-challenge/train_sample_videos/agqphdxmwt.mp4')\nsuccess, frame = v_cap.read()\n#print(success)#True\n#print(frame)# array\nframe = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n#print(frame)\nframe = Image.fromarray(frame) #<PIL.Image.Image image mode=RGB size=1920x1080 at 0x7F15050945C0>\n# from array to PIL\n#print(frame)\nplt.figure(figsize=(12, 8))\nplt.imshow(frame)\nplt.axis('off')\n\n#其中，X变量存储图像，可以是浮点型数组、unit8数组以及PIL图像，如果其为数组，则需满足一下形状：\n   # (1) M*N      此时数组必须为浮点型，其中值为该坐标的灰度；\n   # (2) M*N*3  RGB（浮点型或者unit8类型）\n   # (3) M*N*4  RGBA（浮点型或者unit8类型）\n\n# Detect face\nface = mtcnn(frame)\nface.shape","metadata":{"execution":{"iopub.status.busy":"2022-03-30T07:38:06.637337Z","iopub.execute_input":"2022-03-30T07:38:06.637665Z","iopub.status.idle":"2022-03-30T07:38:11.791972Z","shell.execute_reply.started":"2022-03-30T07:38:06.637613Z","shell.execute_reply":"2022-03-30T07:38:11.791096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a id='3'>Preventing image normalization</a>\n\nBy default, the MTCNN module of `facenet-pytorch` applies fixed image standardization to faces before returning so they are well suited for the package's face recognition model.\n\nIf you want to get out images that look more normal to the human eye, this normalization can be prevented by creating the detector with `post_process=False`.","metadata":{}},{"cell_type":"code","source":"mtcnn = MTCNN(select_largest=False, post_process=True, device='cuda:0')\nface=mtcnn(frame)\n#print(face)  （-1，1）\n#tensor([[[-0.8555, -0.8320, -0.8008,  ..., -0.6523, -0.6992, -0.7383],\n#         [-0.8555, -0.8398, -0.8086,  ..., -0.6992, -0.7305, -0.7539],\n#         [-0.8555, -0.8398, -0.8164,  ..., -0.6445, -0.6758, -0.6914],\n #        ...,\nface1=face.permute(1,2,0)\n#print(face1.shape) torch.Size([160, 160, 3])\nplt.imshow(face1.int().numpy())\n#plt.imshow(face.permute(1, 2, 0).int().numpy())","metadata":{"execution":{"iopub.status.busy":"2022-03-30T07:38:31.018987Z","iopub.execute_input":"2022-03-30T07:38:31.020773Z","iopub.status.idle":"2022-03-30T07:38:31.416320Z","shell.execute_reply.started":"2022-03-30T07:38:31.020704Z","shell.execute_reply":"2022-03-30T07:38:31.415475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create face detector\nmtcnn = MTCNN(select_largest=False, post_process=False, device='cuda:0')\n# post_process {bool} -- Whether or not to post process images tensors before returning.\n           #(default: {True}) 想要imshow成有色彩的，就不要是toTensor\n\n# Detect face\nface = mtcnn(frame)\n#print(face) （0，225）\n#tensor([[[ 18.,  21.,  25.,  ...,  44.,  38.,  33.],\n#         [ 18.,  20.,  24.,  ...,  38.,  34.,  31.],\n#         [ 18.,  20.,  23.,  ...,  45.,  41.,  39.],\n# Visualize\nplt.imshow(face.permute(1, 2, 0).int().numpy())\nplt.axis('off');","metadata":{"execution":{"iopub.status.busy":"2022-03-30T07:38:35.658438Z","iopub.execute_input":"2022-03-30T07:38:35.658761Z","iopub.status.idle":"2022-03-30T07:38:35.886097Z","shell.execute_reply.started":"2022-03-30T07:38:35.658707Z","shell.execute_reply":"2022-03-30T07:38:35.885164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a id='4'>Margin adjustment</a>\n\nDepending on your downstream processing and how fakes can be identified, you may want to add more (or less) of a margin around the detected faces. This is controlled using the `margin` argument.","metadata":{}},{"cell_type":"code","source":"# Create face detector\nmtcnn = MTCNN(margin=40, select_largest=False, post_process=False, device='cuda:0')\n\n#margin {int} -- Margin to add to bounding box, in terms of pixels in the final image. \n #|          Note that the application of the margin differs slightly from the davidsandberg/facenet\n #|          repo, which applies the margin to the original image before resizing, making the margin\n #|          dependent on the original image size (this is a bug in davidsandberg/facenet).\n# |          (default: {0})\n\n\n# Detect face\nface = mtcnn(frame)\n\n# Visualize\nplt.imshow(face.permute(1, 2, 0).int().numpy())\nplt.axis('off');","metadata":{"execution":{"iopub.status.busy":"2022-03-30T07:38:41.187486Z","iopub.execute_input":"2022-03-30T07:38:41.187962Z","iopub.status.idle":"2022-03-30T07:38:41.377891Z","shell.execute_reply.started":"2022-03-30T07:38:41.187890Z","shell.execute_reply":"2022-03-30T07:38:41.376973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a id='5'>Multiple faces in a single image</a>\n\nUsing MTCNN as above will only return a single face from each frame (or None if none are detected). Since some of the videos in the dataset contain more than one face, you will likely want to return all detected faces as any/all of them may have been modified. This is acheived by setting `keep_all=True`","metadata":{}},{"cell_type":"code","source":"# Create face detector\nmtcnn = MTCNN(margin=20, keep_all=True, post_process=False, device='cuda:0')\n\n# keep_all {bool} -- If True, all detected faces are returned, in the order dictated by the\n# |          select_largest parameter. If a save_path is specified, the first face is saved to that\n# |          path and the remaining faces are saved to <save_path>1, <save_path>2 etc.\n\n\n# mtcnn 的输入形式：\n#        - PIL image or list of PIL images\n# |      - numpy.ndarray (uint8) representing either a single image (3D) or a batch of images (4D).\n\n# Load a single image and display\nv_cap = cv2.VideoCapture('/kaggle/input/deepfake-detection-challenge/train_sample_videos/avibnnhwhp.mp4')\nsuccess, frame = v_cap.read()\n# print(frame.shape) (1080, 1920, 3)\nframe = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\nframe = Image.fromarray(frame)\n\nplt.figure(figsize=(12, 8))\nplt.imshow(frame)\nplt.axis('off')\nplt.show()\n\n# Detect face\nfaces = mtcnn(frame)\n# print(faces.shape) torch.Size([2, 3, 160, 160])\n\n# Visualize\nfig, axes = plt.subplots(1, len(faces))\nfor face, ax in zip(faces, axes):\n    print(ax)\n    ax.imshow(face.permute(1, 2, 0).int().numpy())\n    ax.axis('off')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-30T07:38:55.928164Z","iopub.execute_input":"2022-03-30T07:38:55.928479Z","iopub.status.idle":"2022-03-30T07:38:56.925653Z","shell.execute_reply.started":"2022-03-30T07:38:55.928420Z","shell.execute_reply":"2022-03-30T07:38:56.924877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a id='6'>Batched detection</a>\n\n`facenet-pytorch` is also capable of performing face detection on batches of images, typically providing considerable speed-up. A batch should be structured as list of PIL images of equal dimension. The returned object will have an additional first dimension corresponding to the batch. Each image in the batch may have one or more faces detected.\n\nIn the following example, we use MTCNN to detect multiple faces in:\n1. A single batch of frames, and\n1. Every frame of a video","metadata":{}},{"cell_type":"code","source":"help(MTCNN)","metadata":{"execution":{"iopub.status.busy":"2022-03-30T03:48:35.266299Z","iopub.execute_input":"2022-03-30T03:48:35.266596Z","iopub.status.idle":"2022-03-30T03:48:35.276642Z","shell.execute_reply.started":"2022-03-30T03:48:35.266549Z","shell.execute_reply":"2022-03-30T03:48:35.275886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mtcnn = MTCNN(margin=20, keep_all=True, post_process=False, device='cuda:0')\n#post_process {bool} -- Whether or not to post process images tensors before returning.\n #|          (default: {True})\n    \n# Load a video\nv_cap = cv2.VideoCapture('/kaggle/input/deepfake-detection-challenge/train_sample_videos/avibnnhwhp.mp4')\n#print(v_cap) <VideoCapture 0x7f14a0143fb0>\n#print(v_cap.get(cv2.CAP_PROP_FRAME_COUNT)) 300.0\nv_len = int(v_cap.get(cv2.CAP_PROP_FRAME_COUNT))\n\n# Loop through video, taking a handful of frames to form a batch\nframes = []\nfor i in tqdm(range(v_len)):\n    \n    # Load frame\n    success = v_cap.grab()\n    if i % 50 == 0:\n        success, frame = v_cap.retrieve()\n    else:\n        continue\n    if not success:\n        continue\n        \n    # Add to batch\n    frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n    print(frame.shape)\n    frames.append(Image.fromarray(frame))\n\n\n# Detect faces in batch\n#print(np.asarray(frames))\nfaces = mtcnn(frames) #(6,160,160,3)\n#print(faces[0].shape) torch.Size([2, 3, 160, 160])\n\nfig, axes = plt.subplots(len(faces), 2, figsize=(6, 15))\nfor i, frame_faces in enumerate(faces):\n    for j, face in enumerate(frame_faces):\n        axes[i, j].imshow(face.permute(1, 2, 0).int().numpy())\n        axes[i, j].axis('off')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-30T07:45:42.545881Z","iopub.execute_input":"2022-03-30T07:45:42.546218Z","iopub.status.idle":"2022-03-30T07:45:45.716044Z","shell.execute_reply.started":"2022-03-30T07:45:42.546167Z","shell.execute_reply":"2022-03-30T07:45:45.715266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following example uses a similar approach to detect all faces in all frames in a video.","metadata":{}},{"cell_type":"code","source":"# Load a video\nv_cap = cv2.VideoCapture('/kaggle/input/deepfake-detection-challenge/train_sample_videos/avibnnhwhp.mp4')\nv_len = int(v_cap.get(cv2.CAP_PROP_FRAME_COUNT))\n\n# Loop through video\nbatch_size = 16\nframes = []\nfaces = []\n\nfor _ in tqdm(range(v_len)):\n    \n    # Load frame\n    success, frame = v_cap.read()\n    if not success:\n        continue\n        \n    # Add to batch\n    frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n    frames.append(Image.fromarray(frame))\n    \n    # When batch is full, detect faces and reset batch list\n    if len(frames) >= batch_size:\n        faces.extend(mtcnn(frames))\n        frames = []\n\nplt.figure(figsize=(12, 4))\nplt.plot([len(f) for f in faces])\nplt.title('Detected faces per frame');","metadata":{"execution":{"iopub.status.busy":"2022-03-30T07:46:09.843657Z","iopub.execute_input":"2022-03-30T07:46:09.844197Z","iopub.status.idle":"2022-03-30T07:46:28.711880Z","shell.execute_reply.started":"2022-03-30T07:46:09.843940Z","shell.execute_reply":"2022-03-30T07:46:28.711094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a id='7'>Bounding boxes and facial landmarks</a>\n\nTo return bounding boxes and facial landmarks from MTCNN, instead of calling the `mtcnn` object directly, call `mtcnn.detect()` instead.\n\nUnlike the forward method (shown in each of examples above), the `.detect()` method will always return all detected bounding boxes (and optional landmarks) in an image.\n\nThe following example demonstrates the use of the `.detect()` method on a single image.\n\nNote that the `margin` argument, if used when creating the MTCNN detector, is not used in the `detect()` method. `detect()` returns the true bounding boxes, so the margin can be applied subsequently by the user if desired.","metadata":{}},{"cell_type":"markdown","source":"## Create face detector\n#detect(self, img, landmarks=False)\n |      Detect all faces in PIL image and return bounding boxes and optional facial landmarks.\n|      \n |      This method is used by the forward method and is also useful for face detection tasks\n |      that require lower-level handling of bounding boxes and facial landmarks (e.g., face\n |      tracking). The functionality of the forward function can be emulated by using this method\n |      followed by the extract_face() function.\n |      \n# |      Arguments:\n |          img {PIL.Image, np.ndarray, or list} -- A PIL image or a list of PIL images.\n |      \n# |      Keyword Arguments:\n |          landmarks {bool} -- Whether to return facial landmarks in addition to bounding boxes.\n |              (default: {False})\n |      \n# |      Returns:\n |          tuple(numpy.ndarray, list) -- For N detected faces, a tuple containing an\n |              Nx4 array of bounding boxes and a length N list of detection probabilities.\n |              Returned boxes will be sorted in descending order by detection probability if\n |              self.select_largest=False, otherwise the largest face will be returned first.\n |              If `img` is a list of images, the items returned have an extra dimension\n |              (batch) as the first dimension. Optionally, a third item, the facial landmarks,\n |              are returned if `landmarks=True`.\n |      \n |      Example:\n |      >>> from PIL import Image, ImageDraw\n |      >>> from facenet_pytorch import MTCNN, extract_face\n |      >>> mtcnn = MTCNN(keep_all=True)\n |      >>> boxes, probs, points = mtcnn.detect(img, landmarks=True)\n |      >>> # Draw boxes and save faces\n |      >>> img_draw = img.copy()\n |      >>> draw = ImageDraw.Draw(img_draw)\n |      >>> for i, (box, point) in enumerate(zip(boxes, points)):\n |      ...     draw.rectangle(box.tolist(), width=5)\n |      ...     for p in point:\n |      ...         draw.rectangle((p - 10).tolist() + (p + 10).tolist(), width=10)\n |      ...     extract_face(img, box, save_path='detected_face_{}.png'.format(i))\n |      >>> img_draw.save('annotated_faces.png')\n |  \n","metadata":{}},{"cell_type":"code","source":"# Create face detector\nmtcnn = MTCNN(keep_all=True, device='cuda:0')\n\n# Load a single image and display\nv_cap = cv2.VideoCapture('/kaggle/input/deepfake-detection-challenge/train_sample_videos/agqphdxmwt.mp4')\nsuccess, frame = v_cap.read()\nframe = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n#print(frame.shape)\nframe = Image.fromarray(frame)\n\n# Detect face\nboxes, probs, landmarks = mtcnn.detect(frame, landmarks=True,)\n#脸最大的\n\n#print(boxes) [[1037.7498    52.9199  1143.4684   193.87721]]\n#网络的第二部分给出框的精确位置，一般称为框回归。P-Net输入的12×12的图像块可能并不是完美的人脸框的位置，如有的时候人脸并不正好为方形，有可能12×12的图像偏左或偏右，因此需要输出当前框位置相对完美的人脸框位置的偏移。\n#这个偏移大小为1×1×4，即表示框左上角的横坐标的相对偏移，框左上角的纵坐标的相对偏移、框的宽度的误差、框的高度的误差。\n\n#print(probs) [0.99999654]\n#网络的第一部分输出是用来判断该图像是否包含人脸，输出向量大小为1×1×2，也就是两个值，即图像是人脸的概率和图像不是人脸的概率\n\n#print(landmarks)\n#5个关键点分别对应着左眼的位置、右眼的位置、鼻子的位置、左嘴巴的位置、右嘴巴的位置。每个关键点需要两维来表示，因此输出是向量大小为1×1×10\n#[[[1071.0398    106.883896]\n#  [1119.9302    108.41072 ]\n#  [1097.3862    131.19383 ]\n#  [1072.6781    158.70648 ]\n#  [1117.6167    159.1854  ]]]\n# Visualize\nfig, ax = plt.subplots(figsize=(16, 12))\nax.imshow(frame)\nax.axis('off')\n\nfor box, landmark in zip(boxes, landmarks):\n    ax.scatter(*np.meshgrid(box[[0, 2]], box[[1, 3]]))\n    ax.scatter(landmark[:, 0], landmark[:, 1],s=8)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-30T07:57:26.224171Z","iopub.execute_input":"2022-03-30T07:57:26.224506Z","iopub.status.idle":"2022-03-30T07:57:26.880211Z","shell.execute_reply.started":"2022-03-30T07:57:26.224447Z","shell.execute_reply":"2022-03-30T07:57:26.879491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(16, 12))\nax.imshow(frame)\nax.axis('off')\nfor box, landmark in zip(boxes, landmarks):\n    print(box)\n    print(landmark)\n    ax.scatter(*np.meshgrid(box[[0, 2]], box[[1, 3]]))\n    ax.scatter(landmark[:, 0], landmark[:, 1])","metadata":{"execution":{"iopub.status.busy":"2022-03-30T08:01:21.599225Z","iopub.execute_input":"2022-03-30T08:01:21.599535Z","iopub.status.idle":"2022-03-30T08:01:22.135264Z","shell.execute_reply.started":"2022-03-30T08:01:21.599482Z","shell.execute_reply":"2022-03-30T08:01:22.134482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following example demonstrates how to show bounding boxes and facial landmarks in every frame in a video.","metadata":{}},{"cell_type":"code","source":"v_cap = cv2.VideoCapture('/kaggle/input/deepfake-detection-challenge/train_sample_videos/avibnnhwhp.mp4')\n\n# Loop through video\nbatch_size = 32\nframes = []\nboxes = []\nlandmarks = []\nview_frames = []\nview_boxes = []\nview_landmarks = []\nfor _ in tqdm(range(v_len)):\n    \n    # Load frame\n    success, frame = v_cap.read()\n    if not success:\n        continue\n        \n    # Add to batch, resizing for speed\n    frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n    frame = Image.fromarray(frame)\n    frame = frame.resize([int(f * 0.25) for f in frame.size])# 0，0.25，0.5...40\n    \n    frames.append(frame)\n\n    if len(frames) >= batch_size:\n        batch_boxes, _, batch_landmarks = mtcnn.detect(frames, landmarks=True)\n        boxes.extend(batch_boxes)\n        landmarks.extend(batch_landmarks)\n        \n        view_frames.append(frames[-1])# 第31个\n        view_boxes.append(boxes[-1])\n        view_landmarks.append(landmarks[-1])\n        \n        #frames = []\n        \n        break\n    ","metadata":{"execution":{"iopub.status.busy":"2022-03-30T08:13:01.329297Z","iopub.execute_input":"2022-03-30T08:13:01.329749Z","iopub.status.idle":"2022-03-30T08:13:02.218420Z","shell.execute_reply.started":"2022-03-30T08:13:01.329536Z","shell.execute_reply":"2022-03-30T08:13:02.217625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(frames[23].size)\nprint(len(frames))","metadata":{"execution":{"iopub.status.busy":"2022-03-30T08:15:03.940310Z","iopub.execute_input":"2022-03-30T08:15:03.940642Z","iopub.status.idle":"2022-03-30T08:15:03.945881Z","shell.execute_reply.started":"2022-03-30T08:15:03.940585Z","shell.execute_reply":"2022-03-30T08:15:03.944733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"batch_boxes, _, batch_landmarks = mtcnn.detect(frames, landmarks=True)","metadata":{"execution":{"iopub.status.busy":"2022-03-30T08:15:15.435529Z","iopub.execute_input":"2022-03-30T08:15:15.435863Z","iopub.status.idle":"2022-03-30T08:15:15.872198Z","shell.execute_reply.started":"2022-03-30T08:15:15.435810Z","shell.execute_reply":"2022-03-30T08:15:15.871292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(batch_boxes.shape)\nprint( batch_landmarks.shape)","metadata":{"execution":{"iopub.status.busy":"2022-03-30T08:16:09.316075Z","iopub.execute_input":"2022-03-30T08:16:09.316374Z","iopub.status.idle":"2022-03-30T08:16:09.322207Z","shell.execute_reply.started":"2022-03-30T08:16:09.316325Z","shell.execute_reply":"2022-03-30T08:16:09.321273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"frames[-1]","metadata":{"execution":{"iopub.status.busy":"2022-03-30T08:17:08.230052Z","iopub.execute_input":"2022-03-30T08:17:08.230345Z","iopub.status.idle":"2022-03-30T08:17:08.298228Z","shell.execute_reply.started":"2022-03-30T08:17:08.230294Z","shell.execute_reply":"2022-03-30T08:17:08.297494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"boxes[-1]","metadata":{"execution":{"iopub.status.busy":"2022-03-30T08:17:57.205424Z","iopub.execute_input":"2022-03-30T08:17:57.205793Z","iopub.status.idle":"2022-03-30T08:17:57.212748Z","shell.execute_reply.started":"2022-03-30T08:17:57.205736Z","shell.execute_reply":"2022-03-30T08:17:57.211931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"boxes.extend(batch_boxes)\nlandmarks.extend(batch_landmarks)\n        \nview_frames.append(frames[-1])\nview_boxes.append(boxes[-1])\nview_landmarks.append(landmarks[-1])","metadata":{"execution":{"iopub.status.busy":"2022-03-30T08:18:42.767315Z","iopub.execute_input":"2022-03-30T08:18:42.767633Z","iopub.status.idle":"2022-03-30T08:18:42.772250Z","shell.execute_reply.started":"2022-03-30T08:18:42.767577Z","shell.execute_reply":"2022-03-30T08:18:42.771489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(view_frames)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"view_frames.append(frames[-1])\nview_boxes.append(boxes[-1])\nview_landmarks.append(landmarks[-1])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load a video\nv_cap = cv2.VideoCapture('/kaggle/input/deepfake-detection-challenge/train_sample_videos/avibnnhwhp.mp4')\n\n# Loop through video\nbatch_size = 32\nframes = []\nboxes = []\nlandmarks = []\nview_frames = []\nview_boxes = []\nview_landmarks = []\nfor _ in tqdm(range(v_len)):\n    \n    # Load frame\n    success, frame = v_cap.read()\n    if not success:\n        continue\n        \n    # Add to batch, resizing for speed\n    frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n    frame = Image.fromarray(frame)\n    frame = frame.resize([int(f * 0.25) for f in frame.size])# 0，0.25，0.5...40\n    frames.append(frame)\n    \n    # When batch is full, detect faces and reset batch list\n    if len(frames) >= batch_size:\n        batch_boxes, _, batch_landmarks = mtcnn.detect(frames, landmarks=True)\n        boxes.extend(batch_boxes)\n        landmarks.extend(batch_landmarks)\n        \n        view_frames.append(frames[-1])\n        view_boxes.append(boxes[-1])\n        view_landmarks.append(landmarks[-1])\n        \n        frames = []\n\n# Visualize\nfig, ax = plt.subplots(3, 3, figsize=(18, 12))\nfor i in range(9): #300/32=9\n    ax[int(i / 3), i % 3].imshow(view_frames[i])\n    ax[int(i / 3), i % 3].axis('off')\n    for box, landmark in zip(view_boxes[i], view_landmarks[i]):\n        ax[int(i / 3), i % 3].scatter(*np.meshgrid(box[[0, 2]], box[[1, 3]]), s=8)\n        ax[int(i / 3), i % 3].scatter(landmark[:, 0], landmark[:, 1], s=6)","metadata":{"execution":{"iopub.status.busy":"2022-03-30T08:31:06.008773Z","iopub.execute_input":"2022-03-30T08:31:06.009080Z","iopub.status.idle":"2022-03-30T08:31:14.989065Z","shell.execute_reply.started":"2022-03-30T08:31:06.009030Z","shell.execute_reply":"2022-03-30T08:31:14.987957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a id='8'>Saving face datasets</a>\n\nIn order to save detected faces directly to file, use MTCNN's `save_path` argument in the forward function. This is compatible with both single images and batch processing.\n\n- For single images, pass a single path string (e.g., '{videoname}\\_{frame}.jpg')\n- For batches of images, pass a list of path strings (one for each frame)\n\nWhen multiple faces are detected in a single image, additional faces are each saved with an incremental integer appended to the end of the save path (e.g., '{videoname}\\_{frame}.jpg' and '{videoname}\\_{frame}\\_1.jpg')\n\nSee example below.","metadata":{}},{"cell_type":"code","source":"help(MTCNN)","metadata":{"execution":{"iopub.status.busy":"2022-03-30T08:32:14.422515Z","iopub.execute_input":"2022-03-30T08:32:14.422854Z","iopub.status.idle":"2022-03-30T08:32:14.438716Z","shell.execute_reply.started":"2022-03-30T08:32:14.422801Z","shell.execute_reply":"2022-03-30T08:32:14.438087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Single image\nmtcnn(frame, save_path='single_image.jpg');\n#forward(self, img, save_path=None, return_prob=False)\n# |      Run MTCNN face detection on a PIL image or numpy array. This method performs both\n# |      detection and extraction of faces, returning tensors representing detected faces rather\n# |      than the bounding boxes. To access bounding boxes, see the MTCNN.detect() method below.\n# |      \n# |      Arguments:\n# |          img {PIL.Image, np.ndarray, or list} -- A PIL image, np.ndarray, or list.\n# |      \n# |      Keyword Arguments:\n# |          save_path {str} -- An optional save path for the cropped image. Note that when\n# |              self.post_process=True, although the returned tensor is post processed, the saved\n# |              face image is not, so it is a true representation of the face in the input image.\n# |              If `img` is a list of images, `save_path` should be a list of equal length.\n# |              (default: {None})\n# Batch\nsave_paths = [f'image_{i}.jpg' for i in range(len(frames))]\nmtcnn(frames, save_path=save_paths);","metadata":{"execution":{"iopub.status.busy":"2022-03-30T08:33:28.102444Z","iopub.execute_input":"2022-03-30T08:33:28.102770Z","iopub.status.idle":"2022-03-30T08:33:28.343215Z","shell.execute_reply.started":"2022-03-30T08:33:28.102716Z","shell.execute_reply":"2022-03-30T08:33:28.342486Z"},"trusted":true},"execution_count":null,"outputs":[]}]}