{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Eight FaceForensics++ deepfake detectors on DFDC videos\n\nThe eight [xdfdet](https://www.kaggle.com/models/mertilovski/xdfdet) models come from the paper *Augmentation and Cutout in Deepfake Detection: A Comparative Study of Accuracy, Calibration, and Attention* (UBMK 2026). They share one EfficientNet-B4 architecture and differ only in data augmentation and cutout. All of them were trained and tested on FaceForensics++ alone.\n\nThis notebook scores them on the 400 labelled sample videos of the [Deepfake Detection Challenge](https://www.kaggle.com/competitions/deepfake-detection-challenge) (DFDC), a dataset none of the models saw during training. DFDC is also the origin of the cutout method studied in the paper: the winning solution by Selim Seferbekov removed facial regions from fake training frames.\n\n[Project page](https://xdfdet.mertkayacs.com) · [Code](https://github.com/mertkayacs/xdfdet) · [Thesis](https://doi.org/10.5281/zenodo.18998566) · [Quickstart notebook](https://www.kaggle.com/code/mertilovski/xdfdet-quickstart)"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"# The released models need the exact versions they were saved with (TensorFlow 2.19, Keras 3.10).\n# pip warns about preinstalled Kaggle packages that expect newer TensorFlow; none of them is used here.\n!pip install -q --progress-bar off git+https://github.com/mertkayacs/xdfdet > /dev/null 2>&1 && echo \"xdfdet installed\""},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"import os\nos.environ[\"TF_CPP_MIN_LOG_LEVEL\"] = \"2\"  # fewer TensorFlow logs; the CUDA \"Unable to register\" lines are harmless\n\nimport glob, json, time\nimport cv2\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nfrom mtcnn import MTCNN\nfrom sklearn.metrics import roc_auc_score\n\nimport xdfdet\nfrom xdfdet.evaluate import metrics\nfrom xdfdet.video import crop_faces, frame_indices, normalize\n\nDATA = \"/kaggle/input/competitions/deepfake-detection-challenge/train_sample_videos\"\nMODELS = \"/kaggle/input/models/mertilovski/xdfdet/keras\"\nprint(xdfdet.__version__, len(glob.glob(f\"{DATA}/*.mp4\")), \"videos\")"},{"cell_type":"markdown","metadata":{},"source":"## Data\n\n`metadata.json` labels each video `REAL` or `FAKE`: 77 real and 323 fake. The classes are unbalanced, so AUC and the per-class accuracies below carry more information than plain accuracy.\n\nFaces go through the same pipeline as in the paper: MTCNN detection, alignment on the eyes, a 224×224 crop from 32 evenly spaced frames, and 12 of those crops (every second one of the first 24) per model input. Two changes keep the run short. MTCNN runs only on the 12 frames the model reads, and it searches a copy of each 1080p frame at half resolution, falling back to full resolution when that finds no face. A video with no face in those 12 frames goes through the standard search over all 32. The crop itself is always cut from the full-resolution frame. On a few test videos these changes moved the video scores by about 0.01."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"labels_raw = json.load(open(f\"{DATA}/metadata.json\"))\nvideos = sorted(labels_raw)\ny = np.array([labels_raw[v][\"label\"] == \"REAL\" for v in videos], dtype=int)  # 1 = real, as in the paper\nprint(f\"{y.sum()} real, {len(y) - y.sum()} fake\")\n\n\nclass FastMTCNN:\n    \"\"\"MTCNN for `crop_faces`, run only on the crops the model reads, at half resolution.\n\n    `crop_faces` calls `detect_faces` once per sampled frame, in order. For frames the\n    model never sees it returns no face, so `crop_faces` reuses the previous box.\n    \"\"\"\n\n    def __init__(self, detector, scale=0.5):\n        self.detector, self.scale, self.calls = detector, scale, 0\n        self.needed = set(frame_indices(32))\n\n    def detect_faces(self, frame):\n        k, self.calls = self.calls, self.calls + 1\n        if k not in self.needed:\n            return []\n        small = cv2.resize(frame, None, fx=self.scale, fy=self.scale, interpolation=cv2.INTER_AREA)\n        faces = self.detector.detect_faces(small)\n        if not faces:  # small faces can vanish at half size; search the full frame instead\n            return self.detector.detect_faces(frame)\n        for f in faces:\n            f[\"box\"] = [round(v / self.scale) for v in f[\"box\"]]\n            f[\"keypoints\"] = {k: tuple(round(c / self.scale) for c in p) for k, p in f[\"keypoints\"].items()}\n        return faces"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"mtcnn = MTCNN()\nframes, missing = {}, []\nstart = time.time()\nfor i, v in enumerate(videos):\n    try:\n        crops = crop_faces(f\"{DATA}/{v}\", detector=FastMTCNN(mtcnn))\n    except ValueError:\n        try:  # no face in the 12 frames: fall back to the full search over all 32\n            crops = crop_faces(f\"{DATA}/{v}\", detector=mtcnn)\n        except ValueError:\n            missing.append(v)\n            continue\n    frames[v] = np.stack([crops[j] for j in frame_indices(len(crops))])  # (12, 224, 224, 3) uint8\n    if (i + 1) % 50 == 0:\n        print(f\"{i + 1}/{len(videos)} videos, {(time.time() - start) / 60:.0f} min\")\nprint(len(frames), \"videos cropped; no face found in\", missing)"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"fig, axes = plt.subplots(2, 6, figsize=(12, 4.4))\nexamples = [v for v in videos if v in frames and labels_raw[v][\"label\"] == \"REAL\"][:6]\nfor ax, v in zip(axes[0], examples):\n    ax.imshow(frames[v][6]); ax.set_title(f\"{v[:6]} real\", fontsize=9)\nfakes = [v for v in videos if v in frames and labels_raw[v].get(\"original\") in examples][:6]\nfor ax, v in zip(axes[1], fakes):\n    ax.imshow(frames[v][6]); ax.set_title(f\"{v[:6]} fake\", fontsize=9)\nfor ax in axes.flat:\n    ax.axis(\"off\")\nplt.suptitle(\"Aligned face crops (frame 7 of 12)\", fontsize=11)\nplt.tight_layout()"},{"cell_type":"markdown","metadata":{},"source":"## Scoring\n\nEach model outputs the probability that a frame is real. The video score is the mean over the 12 frames, and a score below 0.5 means fake. The models are loaded one at a time to keep GPU memory free."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"kept = [v for v in videos if v in frames]\ny_kept = np.array([labels_raw[v][\"label\"] == \"REAL\" for v in kept], dtype=int)\n\n\ndef batches(size=16):\n    \"\"\"Normalized input in small batches, so only the uint8 crops stay in memory.\"\"\"\n    for s in range(0, len(kept), size):\n        yield np.stack([np.stack([normalize(f) for f in frames[v]]) for v in kept[s:s + size]])\n\n\n# All eight share one architecture, so one model object is built and each checkpoint's\n# weights are loaded into it. Loading eight full models in a row exhausts the session's RAM.\nscores, model = {}, None\nfor name in xdfdet.RELEASED:\n    path = glob.glob(f\"{MODELS}/{name}/*/{name}.keras\")[0]\n    if model is None:\n        model = xdfdet.load_model(path)\n    else:\n        model.load_weights(path)\n    scores[name] = np.concatenate([np.asarray(model.predict_on_batch(x)).mean(axis=(1, 2)) for x in batches()])\n    pd.DataFrame({\"video\": kept, \"real\": y_kept, **scores}).to_csv(\"dfdc_scores.csv\", index=False)\n    print(name, \"done\")"},{"cell_type":"markdown","metadata":{},"source":"## Results\n\nThe FaceForensics++ AUC is each checkpoint's own test result from the paper's runs. The DFDC columns use the same metric definitions as the paper (`xdfdet.evaluate.metrics`, where F1 treats \"real\" as the positive class), plus the share of real and of fake videos classified correctly at the 0.5 threshold."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"FFPP_AUC = {\"aug-cutout-black\": 0.8981, \"aug-cutout-random\": 0.8820, \"aug-cutout-white\": 0.8734,\n            \"cutout-white\": 0.8700, \"baseline\": 0.8684, \"cutout-black\": 0.8669,\n            \"cutout-random\": 0.8642, \"aug-standard\": 0.8616}\n\nrows = []\nfor name, s in scores.items():\n    m = metrics(y_kept, s)\n    pred_real = s >= 0.5\n    rows.append({\"model\": name, \"FF++ AUC\": FFPP_AUC[name], \"DFDC AUC\": m[\"auc\"], \"F1\": m[\"f1\"],\n                 \"Brier\": m[\"brier\"], \"LogLoss\": m[\"logloss\"],\n                 \"real correct\": pred_real[y_kept == 1].mean(), \"fake correct\": (~pred_real)[y_kept == 0].mean()})\ntable = pd.DataFrame(rows).sort_values(\"DFDC AUC\", ascending=False).set_index(\"model\")\ntable.round(3)"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"# Bootstrap 95% interval of the DFDC AUC, resampling videos.\nrng = np.random.default_rng(0)\nidx = [rng.integers(0, len(y_kept), len(y_kept)) for _ in range(2000)]\nidx = [i for i in idx if 0 < y_kept[i].sum() < len(i)]\nci = {n: np.percentile([roc_auc_score(y_kept[i], s[i]) for i in idx], [2.5, 97.5]) for n, s in scores.items()}\npd.DataFrame(ci, index=[\"low\", \"high\"]).T.loc[table.index].round(3)"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"order = table.index[::-1]\nfig, ax = plt.subplots(figsize=(8, 4.2))\nfor i, name in enumerate(order):\n    a, b = table.loc[name, \"FF++ AUC\"], table.loc[name, \"DFDC AUC\"]\n    ax.plot([b, a], [i, i], color=\"#c9c2b4\", lw=2, zorder=1)\n    ax.plot([ci[name][0], ci[name][1]], [i - 0.18, i - 0.18], color=\"#7a1f2b\", lw=1, alpha=0.6)\nax.scatter(table.loc[order, \"FF++ AUC\"], range(len(order)), s=60, color=\"#1f2a44\", label=\"FaceForensics++ test\", zorder=2)\nax.scatter(table.loc[order, \"DFDC AUC\"], range(len(order)), s=60, color=\"#7a1f2b\", label=\"DFDC sample (line: 95% CI)\", zorder=2)\nax.axvline(0.5, color=\"#888\", lw=1, ls=\":\")\nax.text(0.502, len(order) - 0.4, \"chance\", fontsize=8, color=\"#666\")\nax.set_yticks(range(len(order)), order)\nax.set_xlabel(\"AUC\")\nax.set_xlim(0.3, 1.0)\nax.spines[[\"top\", \"right\"]].set_visible(False)\nax.legend(frameon=False, loc=\"lower left\", fontsize=9)\nax.set_title(\"Same checkpoints, two datasets\", fontsize=11, loc=\"left\")\nplt.tight_layout()"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"fig, axes = plt.subplots(1, 2, figsize=(10, 3.4), sharey=True)\nbins = np.linspace(0, 1, 21)\nfor ax, name in zip(axes, [\"aug-cutout-black\", \"baseline\"]):\n    s = scores[name]\n    ax.hist(s[y_kept == 1], bins=bins, density=True, alpha=0.75, color=\"#1f2a44\", label=\"real videos\")\n    ax.hist(s[y_kept == 0], bins=bins, density=True, alpha=0.55, color=\"#b8893a\", label=\"fake videos\")\n    ax.axvline(0.5, color=\"#888\", lw=1, ls=\":\")\n    ax.set_title(name, fontsize=10, loc=\"left\")\n    ax.set_xlabel(\"predicted real-probability\")\n    ax.spines[[\"top\", \"right\"]].set_visible(False)\naxes[0].set_ylabel(\"density\")\naxes[0].legend(frameon=False, fontsize=9)\nplt.tight_layout()"},{"cell_type":"markdown","metadata":{},"source":"## Findings\n\n**AUC drops, and the ranking barely carries over.** On DFDC the eight models reach an AUC of 0.60 to 0.66, against 0.86 to 0.90 on FaceForensics++. Every 95% interval stays above 0.5, so each model keeps some signal. The intervals also overlap almost completely, and 398 videos cannot separate the models: `aug-cutout-black` has the highest DFDC AUC (0.661) with `baseline` close behind (0.658). The two datasets rank the eight configurations in weakly related orders (Spearman ρ = 0.19).\n\n**Most fakes pass as real.** At the 0.5 threshold every model labels most videos real. The models recognise 79% to 99% of the real videos and catch 15% to 36% of the fakes.\n\n**Calibration degrades the most.** On FaceForensics++ the Brier scores lie between 0.12 and 0.16. On DFDC they rise to 0.44 to 0.65, with LogLoss between 1.8 and 4.9. `aug-cutout-black`, which had the best Brier score on FaceForensics++, is the most overconfident model here: the fake videos receive a mean real-probability of 0.83 from it. The baseline is the best calibrated of the eight on DFDC.\n\n### Reading the result\n\nThe FaceForensics++ training videos come from YouTube, each with a trackable, mostly frontal face, and the fakes used here come from four manipulation methods. DFDC was recorded with 3,426 paid actors, the sample videos are 1080p, and its fakes come from several Deepfake, GAN-based and non-learned methods. Augmentation and cutout improved the detectors on the distribution they were trained on and did not prepare them for this shift; the best in-domain configuration became the most confident one out of domain. Accuracy and attention results from a single benchmark therefore describe that benchmark, and a detector meant for real videos needs training data from several sources and an evaluation on held-out datasets.\n\nLimits of this check: the 400 videos are the labelled sample that ships with the competition data, a small fraction of the more than 100,000 DFDC clips; the classes are unbalanced (77 real, 321 fake after face detection), and several fakes share the same source video."},{"cell_type":"markdown","metadata":{},"source":"## Reuse\n\n`dfdc_scores.csv` in the output holds every video's score from all eight models. The [quickstart notebook](https://www.kaggle.com/code/mertilovski/xdfdet-quickstart) shows how to score a single video and draw its Grad-CAM map.\n\nModels: CC BY-NC 4.0. Code: MIT. The DFDC videos are used under the competition's data terms and are not redistributed here."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python"}},"nbformat":4,"nbformat_minor":5}