{"cells":[{"cell_type":"markdown","metadata":{},"source":"# xdfdet quickstart: score a video and see where the detector looks\n\nThis notebook loads two of the eight [xdfdet](https://www.kaggle.com/models/mertilovski/xdfdet) detectors, scores a real video and a deepfake made from it, and draws the Grad-CAM map behind each decision. The models come from the paper *Augmentation and Cutout in Deepfake Detection: A Comparative Study of Accuracy, Calibration, and Attention* (UBMK 2026):\n\n- `aug-cutout-black`: standard augmentation with black-fill cutout, the best configuration in the paper (FaceForensics++ test AUC 0.898)\n- `baseline`: no augmentation and no cutout (AUC 0.868)\n\nBoth are EfficientNet-B4 models that score 12 aligned face crops per video. The video score is the mean probability that the frames are **real**; below 0.5 means fake.\n\n[Project page](https://xdfdet.mertkayacs.com) · [Code](https://github.com/mertkayacs/xdfdet) · [Thesis](https://doi.org/10.5281/zenodo.18998566) · [All eight models on DFDC](https://www.kaggle.com/code/mertilovski/xdfdet-on-dfdc)"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"# The released models need the exact versions they were saved with (TensorFlow 2.19, Keras 3.10).\n# pip warns about preinstalled Kaggle packages that expect newer TensorFlow; none of them is used here.\n!pip install -q --progress-bar off git+https://github.com/mertkayacs/xdfdet > /dev/null 2>&1 && echo \"xdfdet installed\""},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"import os\nos.environ[\"TF_CPP_MIN_LOG_LEVEL\"] = \"2\"  # fewer TensorFlow logs; the CUDA \"Unable to register\" lines are harmless\n\nimport glob, json\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\n\nimport xdfdet\nfrom xdfdet.explain import overlay, region_scores, video_gradcam\nfrom xdfdet.landmarks import landmarks\nfrom xdfdet.video import crop_faces, frame_indices, normalize\n\nDATA = \"/kaggle/input/competitions/deepfake-detection-challenge/train_sample_videos\"\nMODELS = \"/kaggle/input/models/mertilovski/xdfdet/keras\"\nmodels = {name: xdfdet.load_model(glob.glob(f\"{MODELS}/{name}/*/{name}.keras\")[0])\n          for name in [\"aug-cutout-black\", \"baseline\"]}"},{"cell_type":"markdown","metadata":{},"source":"## A real video and its deepfake\n\nThe example comes from the labelled sample of the [Deepfake Detection Challenge](https://www.kaggle.com/competitions/deepfake-detection-challenge). To avoid picking a flattering case, the cell walks through the fake videos in file-name order and keeps the first one whose source video is also in the sample and where dlib finds facial landmarks in both clips (the region chart below needs them).\n\n`crop_faces` runs MTCNN on 32 evenly spaced frames, aligns each face on the eyes and cuts a 224×224 crop. The model receives 12 of them. Each video takes about a minute on a Kaggle GPU session."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"def clip(name):\n    crops = crop_faces(f\"{DATA}/{name}\")\n    return np.stack([crops[i] for i in frame_indices(len(crops))])  # (12, 224, 224, 3) uint8\n\nmeta = json.load(open(f\"{DATA}/metadata.json\"))\nfor fake in sorted(meta):\n    real = meta[fake].get(\"original\")\n    if meta[fake][\"label\"] != \"FAKE\" or real not in meta:\n        continue\n    clips = {\"real\": clip(real), \"fake\": clip(fake)}\n    if all(landmarks(c[len(c) // 2]) is not None for c in clips.values()):\n        break\n    print(\"skipped\", fake, \"(no landmarks)\")\nprint(\"real:\", real, \" fake:\", fake)\n\nfig, axes = plt.subplots(2, 6, figsize=(12, 4.6))\nfor row, label in zip(axes, clips):\n    for ax, k in zip(row, range(0, 12, 2)):\n        ax.imshow(clips[label][k]); ax.axis(\"off\")\n    row[0].set_title(label, loc=\"left\", fontsize=11)\nplt.tight_layout()"},{"cell_type":"markdown","metadata":{},"source":"## Scores\n\n`normalize` applies the ImageNet normalization the models were trained with. Each model returns one probability per frame; the mean is the video score."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"rows = []\nfor label, frames in clips.items():\n    seq = np.stack([normalize(f) for f in frames])\n    for name, model in models.items():\n        frame_scores = model.predict(seq[None], verbose=0)[0, :, 0]\n        score = float(frame_scores.mean())\n        rows.append({\"video\": label, \"model\": name, \"real-probability\": round(score, 3),\n                     \"verdict\": \"real\" if score >= 0.5 else \"fake\",\n                     \"frames judged real\": f\"{(frame_scores >= 0.5).sum()} / 12\"})\nprint(pd.DataFrame(rows).to_string(index=False))"},{"cell_type":"markdown","metadata":{},"source":"## Where the detector looks\n\nGrad-CAM weights the last convolutional layer by the gradient of the predicted class and shows which pixels drove the decision. The map is averaged over frames 1, 5 and 8, as in the paper, and drawn on the middle frame. dlib's 68 facial landmarks then split the face into eight regions, and the bar chart gives the mean activation inside each one (0 to 100)."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"REGION_ORDER = [\"left_eyebrow\", \"right_eyebrow\", \"left_eye\", \"right_eye\", \"nose\", \"outer_mouth\", \"inner_mouth\", \"jaw\"]\nfig, axes = plt.subplots(2, 4, figsize=(13, 6.4), gridspec_kw={\"width_ratios\": [1, 1.3, 1, 1.3]})\nfor r, label in enumerate(clips):\n    frames = clips[label]\n    seq = np.stack([normalize(f) for f in frames])\n    mid = frames[len(frames) // 2]\n    points = landmarks(mid)\n    for c, name in enumerate(models):\n        cam = video_gradcam(models[name], seq)\n        ax_img, ax_bar = axes[r, 2 * c], axes[r, 2 * c + 1]\n        ax_img.imshow(overlay(mid, cam)); ax_img.axis(\"off\")\n        ax_img.set_title(f\"{label} · {name}\", loc=\"left\", fontsize=10)\n        if points is None:\n            ax_bar.text(0.5, 0.5, \"no landmarks found\", ha=\"center\"); ax_bar.axis(\"off\")\n            continue\n        scores = region_scores(cam, points)\n        ax_bar.barh([k.replace(\"_\", \" \") for k in REGION_ORDER], [scores[k] for k in REGION_ORDER], color=\"#1f2a44\", height=0.6)\n        ax_bar.invert_yaxis(); ax_bar.set_xlim(0, 100)\n        ax_bar.spines[[\"top\", \"right\"]].set_visible(False)\n        ax_bar.tick_params(labelsize=8)\nplt.tight_layout()"},{"cell_type":"markdown","metadata":{},"source":"## Reading the result\n\nBoth models call both videos real with a probability close to 1, so the deepfake goes undetected here. The attention maps still differ in the way the paper describes: `aug-cutout-black` concentrates on the eyes and eyebrows in both clips, while the baseline spreads its activation across the whole face of the fake.\n\nA single pair shows how the pipeline works and says little about accuracy. The models were trained and tested on FaceForensics++ only, and DFDC videos differ in resolution, lighting, compression and manipulation method. The [DFDC notebook](https://www.kaggle.com/code/mertilovski/xdfdet-on-dfdc) scores all eight models on the full 400-video sample.\n\n## Your own video\n\nWith internet enabled, the command line tool downloads a model by name and writes the Grad-CAM image:\n\n```bash\nxdfdet predict my_video.mp4 --model aug-cutout-black --gradcam cam.png\n```\n\nModels: CC BY-NC 4.0, for non-commercial research. Code: MIT. The DFDC videos are used under the competition's data terms and are not redistributed here."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python"}},"nbformat":4,"nbformat_minor":5}