{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Baseline submission using Facenet\n\nThis notebook demonstrates how to use the `facenet-pytorch` package to build a rudimentary deepfake detector without training any models. It also demonstrates a method for (1) loading all video frames, (2) finding all faces, and (3) calculating face embeddings at over 30 frames per second (or greater than 1 video per 10 seconds).\n\nThe following steps are performed:\n\n1. Create pretrained facial detection (MTCNN) and recognition (Inception Resnet) models.\n  * See the following kernel for a strided implementation of MTCNN that is able to process all frames in each video: https://www.kaggle.com/timesler/facenet-pytorch-mtcnn-process-every-frame\n  * See the following kernel for a performance comparison for different face detection implementations: https://www.kaggle.com/timesler/comparison-of-face-detection-packages\n1. For each test video, calculate face feature vectors for **ALL** faces in each video.\n1. Calculate the distance from each face to the centroid for its video.\n1. Use these distances as your means of discrimination.\n\nFor (much) better results, finetune the resnet to the fake/real binary classification task instead - this is just a baseline. Alternatively, I'm sure there is much more interesting things that can be done with the feature vectors.","metadata":{}},{"cell_type":"markdown","source":"## Install dependencies","metadata":{}},{"cell_type":"code","source":"!find /kaggle/input -iname \"*.mp4\" | head -5","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T17:08:58.207995Z","iopub.execute_input":"2026-09-05T17:08:58.208235Z","iopub.status.idle":"2026-09-05T17:09:00.141972Z","shell.execute_reply.started":"2026-09-05T17:08:58.208191Z","shell.execute_reply":"2026-09-05T17:09:00.141037Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Install facenet-pytorch\nimport glob\n\n# Recursively locate the dataset folder and wheel file under /kaggle/input,\n# since the exact path/nesting can vary depending on how the dataset was added\ncandidates = glob.glob('/kaggle/input/**/facenet_pytorch-2.2.7-py3-none-any.whl', recursive=True)\nif not candidates:\n    # Fall back to installing from PyPI if the local wheel isn't found\n    print('Local wheel not found under /kaggle/input, installing from PyPI instead')\n    !pip install facenet-pytorch\n    dataset_dir = None\nelse:\n    wheel_path = candidates[0]\n    dataset_dir = wheel_path.rsplit('/', 1)[0]\n    print(f'Found dataset at: {dataset_dir}')\n    !pip install \"$wheel_path\"\n\nfrom facenet_pytorch.models.inception_resnet_v1 import get_torch_home\ntorch_home = get_torch_home()\n\n# Copy model checkpoints to torch cache so they are loaded automatically by the package\n# (only needed if we found the local dataset with pre-downloaded checkpoints)\nif dataset_dir is not None:\n    !mkdir -p $torch_home/checkpoints/\n    !cp \"$dataset_dir/20180402-114759-vggface2-logits.pth\" $torch_home/checkpoints/vggface2_DG3kwML46X.pt\n    !cp \"$dataset_dir/20180402-114759-vggface2-features.pth\" $torch_home/checkpoints/vggface2_G5aNV2VSMn.pt\nelse:\n    print('No local checkpoints found — InceptionResnetV1(pretrained=\\'vggface2\\') will download weights automatically on first use.')","metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"execution":{"iopub.status.busy":"2026-09-05T17:09:00.143850Z","iopub.execute_input":"2026-09-05T17:09:00.144122Z","iopub.status.idle":"2026-09-05T17:09:06.647353Z","shell.execute_reply.started":"2026-09-05T17:09:00.144057Z","shell.execute_reply":"2026-09-05T17:09:06.646488Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Imports","metadata":{}},{"cell_type":"code","source":"import os\nimport glob\nimport time\nimport torch\nimport cv2\nfrom PIL import Image\nimport numpy as np\nimport pandas as pd\nfrom matplotlib import pyplot as plt\nfrom tqdm import tqdm\n\nfrom facenet_pytorch import MTCNN, InceptionResnetV1, extract_face\n\ndevice = 'cuda:0' if torch.cuda.is_available() else 'cpu'\nprint(f'Running on device: {device}')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-09-05T17:09:06.648788Z","iopub.execute_input":"2026-09-05T17:09:06.649030Z","iopub.status.idle":"2026-09-05T17:09:06.654833Z","shell.execute_reply.started":"2026-09-05T17:09:06.648984Z","shell.execute_reply":"2026-09-05T17:09:06.654045Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Create MTCNN and Inception Resnet models\n\nBoth models are pretrained. The Inception Resnet weights will be downloaded the first time it is instantiated; after that, they will be loaded from the torch cache.","metadata":{}},{"cell_type":"code","source":"# Load face detector\nmtcnn = MTCNN(margin=14, keep_all=True, factor=0.5, device=device).eval()\n\n# Load facial recognition model\nresnet = InceptionResnetV1(pretrained='vggface2', device=device).eval()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T17:09:06.655969Z","iopub.execute_input":"2026-09-05T17:09:06.656226Z","iopub.status.idle":"2026-09-05T17:09:07.368124Z","shell.execute_reply.started":"2026-09-05T17:09:06.656168Z","shell.execute_reply":"2026-09-05T17:09:07.367125Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Process test videos\n\nAfter defining a few helper functions, this code loops through all videos and passes **_all_** frames from each through the face detector followed by facenet. Finally, we calculate the distance from the centroid to the extracted feature for each face.","metadata":{}},{"cell_type":"code","source":"   import os\n   if os.path.exists('checkpoint.pkl'):\n       os.remove('checkpoint.pkl')\n\n   !ls /kaggle/input/","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T17:09:07.370654Z","iopub.execute_input":"2026-09-05T17:09:07.370819Z","iopub.status.idle":"2026-09-05T17:09:08.144181Z","shell.execute_reply.started":"2026-09-05T17:09:07.370791Z","shell.execute_reply":"2026-09-05T17:09:08.143238Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class DetectionPipeline:\n    \"\"\"Pipeline class for detecting faces in the frames of a video file.\"\"\"\n    \n    def __init__(self, detector, n_frames=None, batch_size=60, resize=None):\n        \"\"\"Constructor for DetectionPipeline class.\n        \n        Keyword Arguments:\n            n_frames {int} -- Total number of frames to load. These will be evenly spaced\n                throughout the video. If not specified (i.e., None), all frames will be loaded.\n                (default: {None})\n            batch_size {int} -- Batch size to use with MTCNN face detector. (default: {32})\n            resize {float} -- Fraction by which to resize frames from original prior to face\n                detection. A value less than 1 results in downsampling and a value greater than\n                1 result in upsampling. (default: {None})\n        \"\"\"\n        self.detector = detector\n        self.n_frames = n_frames\n        self.batch_size = batch_size\n        self.resize = resize\n    \n    def __call__(self, filename):\n        \"\"\"Load frames from an MP4 video and detect faces.\n\n        Arguments:\n            filename {str} -- Path to video.\n        \"\"\"\n        # Create video reader and find length\n        v_cap = cv2.VideoCapture(filename)\n        v_len = int(v_cap.get(cv2.CAP_PROP_FRAME_COUNT))\n\n        # Pick 'n_frames' evenly spaced frames to sample\n        if self.n_frames is None:\n            sample = np.arange(0, v_len)\n        else:\n            sample = np.linspace(0, v_len - 1, self.n_frames).astype(int)\n\n        # Loop through frames\n        faces = []\n        frames = []\n        for j in range(v_len):\n            success = v_cap.grab()\n            if j in sample:\n                # Load frame\n                success, frame = v_cap.retrieve()\n                if not success:\n                    continue\n                frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n                frame = Image.fromarray(frame)\n                \n                # Resize frame to desired size\n                if self.resize is not None:\n                    frame = frame.resize([int(d * self.resize) for d in frame.size])\n                frames.append(frame)\n\n                # When batch is full, detect faces and reset frame list\n                if len(frames) % self.batch_size == 0 or j == sample[-1]:\n                    faces.extend(self.detector(frames))\n                    frames = []\n\n        v_cap.release()\n\n        return faces    \n\n\ndef process_faces(faces, resnet):\n    # Filter out frames without faces\n    faces = [f for f in faces if f is not None]\n    faces = torch.cat(faces).to(device)\n\n    # Generate facial feature vectors using a pretrained model\n    embeddings = resnet(faces)\n\n    # Calculate centroid for video and distance of each face's feature vector from centroid\n    centroid = embeddings.mean(dim=0)\n    x = (embeddings - centroid).norm(dim=1).cpu().numpy()\n    \n    return x","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T17:09:08.145899Z","iopub.execute_input":"2026-09-05T17:09:08.146158Z","iopub.status.idle":"2026-09-05T17:09:08.157774Z","shell.execute_reply.started":"2026-09-05T17:09:08.146106Z","shell.execute_reply":"2026-09-05T17:09:08.156744Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define face detection pipeline\ndetection_pipeline = DetectionPipeline(detector=mtcnn, batch_size=16, resize=0.25)\n\n# Get all test videos\nfilenames = glob.glob('/kaggle/input/competitions/deepfake-detection-challenge/test_videos/*.mp4')\nfilenames = filenames[:20]\n# --- Checkpoint setup ---\nCHECKPOINT_PATH = 'checkpoint.pkl'\nCHECKPOINT_EVERY = 25  # save progress every N videos\n\nimport pickle\n\ndef save_checkpoint(path, filenames, X, processed_idx):\n    with open(path, 'wb') as f:\n        pickle.dump({'filenames': filenames, 'X': X, 'processed_idx': processed_idx}, f)\n\ndef load_checkpoint(path):\n    if os.path.exists(path):\n        with open(path, 'rb') as f:\n            return pickle.load(f)\n    return None\n\nckpt = load_checkpoint(CHECKPOINT_PATH)\nif ckpt is not None and ckpt['filenames'] == filenames:\n    X = ckpt['X']\n    start_idx = ckpt['processed_idx'] + 1\n    print(f'Resuming from checkpoint at video {start_idx}/{len(filenames)}')\nelse:\n    X = []\n    start_idx = 0\n\nstart = time.time()\nn_processed = 0\nwith torch.no_grad():\n    for i, filename in enumerate(tqdm(filenames)):\n        if i < start_idx:\n            continue\n        try:\n            # Load frames and find faces\n            faces = detection_pipeline(filename)\n\n            # Calculate embeddings\n            X.append(process_faces(faces, resnet))\n\n        except KeyboardInterrupt:\n            print('\\nStopped.')\n            save_checkpoint(CHECKPOINT_PATH, filenames, X, i - 1)\n            break\n\n        except Exception as e:\n            print(e)\n            X.append(None)\n\n        n_processed += len(faces)\n        print(f'Frames per second (load+detect+embed): {n_processed / (time.time() - start):6.3}\\r', end='')\n\n        if (i + 1) % CHECKPOINT_EVERY == 0:\n            save_checkpoint(CHECKPOINT_PATH, filenames, X, i)\n\n    else:\n        # Loop completed without break: save final checkpoint too\n        save_checkpoint(CHECKPOINT_PATH, filenames, X, len(filenames) - 1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T17:09:08.158981Z","iopub.execute_input":"2026-09-05T17:09:08.159166Z","iopub.status.idle":"2026-09-05T17:16:09.070394Z","shell.execute_reply.started":"2026-09-05T17:09:08.159135Z","shell.execute_reply":"2026-09-05T17:16:09.069730Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Predict classes\n\nThe below weights were selected by following the same process as above for the train sample videos and then using a logistic regression model to fit to the labels. Note that, intuitively, this is not a very good approach as it does nothing to take into account the progression of feature vectors throughout a video, just combines them together using the weights below. This step is provided as a placeholder only; it should be replaced with a more thoughtful mapping from a sequence of feature vectors to a single prediction.","metadata":{}},{"cell_type":"code","source":"bias = -0.2942\nweight = 0.68235746\n\nsubmission = []\nfor filename, x_i in zip(filenames, X):\n    if x_i is not None:\n        prob = 1 / (1 + np.exp(-(bias + (weight * x_i).mean())))\n    else:\n        prob = 0.5\n    submission.append([os.path.basename(filename), prob])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T17:16:09.071273Z","iopub.execute_input":"2026-09-05T17:16:09.071482Z","iopub.status.idle":"2026-09-05T17:16:09.076410Z","shell.execute_reply.started":"2026-09-05T17:16:09.071444Z","shell.execute_reply":"2026-09-05T17:16:09.075764Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Visualize results","metadata":{}},{"cell_type":"code","source":"# Distribution of per-video mean distance-from-centroid (raw signal before calibration)\nmean_distances = [x_i.mean() if x_i is not None else np.nan for x_i in X]\nmean_distances = np.array(mean_distances)\n\nfig, axes = plt.subplots(1, 3, figsize=(18, 5))\n\n# 1. Histogram of mean distance-from-centroid per video\naxes[0].hist(mean_distances[~np.isnan(mean_distances)], bins=30, color='steelblue', edgecolor='black')\naxes[0].set_title('Mean Distance from Centroid per Video')\naxes[0].set_xlabel('Mean distance')\naxes[0].set_ylabel('Number of videos')\n\n# 2. Histogram of predicted probabilities\nprobs = [row[1] for row in submission]\naxes[1].hist(probs, bins=30, color='indianred', edgecolor='black')\naxes[1].set_title('Predicted Fake Probability Distribution')\naxes[1].set_xlabel('Predicted probability')\naxes[1].set_ylabel('Number of videos')\n\n# 3. Number of faces detected per video (processing coverage)\nn_faces_per_video = [len(x_i) if x_i is not None else 0 for x_i in X]\naxes[2].hist(n_faces_per_video, bins=30, color='seagreen', edgecolor='black')\naxes[2].set_title('Faces Detected per Video')\naxes[2].set_xlabel('Number of faces')\naxes[2].set_ylabel('Number of videos')\n\nplt.tight_layout()\nplt.show()\n\n# Summary stats\nn_failed = sum(1 for x_i in X if x_i is None)\nprint(f'Videos processed: {len(X)}')\nprint(f'Videos with no usable faces/errors: {n_failed}')\nprint(f'Mean predicted probability: {np.nanmean(probs):.4f}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-05T17:16:09.077255Z","iopub.execute_input":"2026-09-05T17:16:09.077817Z","iopub.status.idle":"2026-09-05T17:16:09.718566Z","shell.execute_reply.started":"2026-09-05T17:16:09.077440Z","shell.execute_reply":"2026-09-05T17:16:09.717946Z"}},"outputs":[],"execution_count":null}]}