Detection, tracking & alignment
Find the face, then hold onto it
A reliable swap depends on three stacked systems: a detector that finds faces, a landmarker that aligns and tracks them frame to frame, and a recognizer that matches them to an identity. Each has its own model choice and a score threshold you tune for recall vs. precision.
Detector
The detector localizes faces in each frame. Choice of model, resolution and angles controls how well it handles small, tilted, or moving faces.
| Model | VRAM | Why / when |
|---|---|---|
| retinaface | mid | Robust multi-scale face detector; historically the default. Balanced default detection. |
| SCRFDrecent | low | Sample-and-Computation-Redistribution detector - efficient with strong accuracy. Speed + accuracy on real hardware. |
| yoloface | low | Light YOLO-based detector - fastest, lowest cost. Low-VRAM and fast scans. |
--face-detector-size320 - 640
Larger = higher recall & detail cost. Typical default ~512 (version-dependent).
--face-detector-score0 - 1
Lower accepts weaker detections (more faces, more false positives).
--face-detector-angles0 / 90 / 180 / 270
Facefusion rotates the frame internally to catch tilted faces. Adding 90/180/270 raises recall on head-tilted shots at a throughput cost.
--face-detector-model: Newer builds also expose a 'many' choice that runs several detectors and merges results - the most robust and the slowest.
Landmarker
After a box is found, the landmarker defines the facial geometry used for alignment and tracking. A 68-point landmarker gives denser, more stable alignment than a lightweight 5-point option.
| Model | VRAM | Why / when |
|---|---|---|
| 2dfan4 | mid | 68-point landmarker (the default) - gives dense landmarks across the whole face for stable alignment. Default; best for smooth tracking and alignment. |
| pipnetrecent | low | A lighter Pose-Invariant landmarker added in newer builds for fast 5-point / pose alignment. Fast runs and constrained hardware. |
--face-landmarker-score0 - 1
Lower = more permissive landmark acceptance, useful for extreme angles.
--face-landmarker-model2dfan4 / pipnet
pipnet is added in newer builds.
Recognition
ArcFace-style embedding (arcface_w600k_r50) turns a face into a 512-dim vector. Facefusion compares embeddings with a distance metric and a threshold to decide whether two faces match - used to restrict which target faces get swapped and to verify identity.
--face-recognizer-model- arcface_w600k_r50, arcface_w600k_r50_v2.--face-recognizer-distance- cosine_distance, euclidean_distance.--face-recognizer-metric- max, min, average.--face-recognizer-threshold- Cosine default ~0.4 (version-dependent); lower = stricter match, higher = looser. Euclidean uses a different scale/default.
Cosine distance cutoff
Facefusion embeds faces into a vector space and treats two faces as the same identity when their distance falls under the threshold. A lower threshold is stricter (fewer faces match, fewer false positives); a higher one is looser. For cosine distance the default sits around 0.4 but the exact value is version-dependent.
Detection example
ff-run \
--face-detector-model scrfd \
--face-detector-size 640x640 \
--face-detector-score 0.4 \
--face-detector-angles 0,90,180,270 \
--face-landmarker-model 2dfan4 \
--face-landmarker-score 0.4 \
--processors face_debugger \
--source <source> --target <target> --output <out>Lowering the detector and landmarker scores raises recall on hard angles but also admits weak detections. Preview with --processors face_debugger before committing to a full render.