Architecture & execution
Where the models actually run
Facefusion is a model orchestration layer: it parses your flags, loads ONNX models into a chosen execution provider, and shuttles frames through them. The three architectural levers are which provider runs the models, how many worker threads/queued frames sit between stages, and how much VRAM each model may claim.
Execution providers
execution_providers.pyEach provider maps the model graphs onto different hardware backends. The same models run everywhere - the provider only decides where and how fast.
| Provider | Platform | VRAM | Notes |
|---|---|---|---|
| cudadefault | NVIDIA GPU | high | The most-tested path. Uses NVIDIA's toolkit through the executor, with cudnn kernels for conv/fusion ops. Best raw throughput on NVIDIA hardware. OnnxRuntime CUDA + TensorRT are separate; CUDA is the baseline NVIDIA provider. Flag takes the form cuda. |
| tensorrt | NVIDIA GPU (tuned) | high | NVIDIA's inference optimizer. First run builds an engine cache (slower warm-up), later runs are often significantly faster at higher throughput than plain CUDA. Engine cache and builder flags are version-dependent (--tensorrt-* args). Best reserved for repeat runs of the same model. |
| directml | AMD / Intel / NVIDIA (Windows) | mid | Microsoft's hardware-agnostic GPU API. Common fallback on Windows laptops (especially AMD iGPUs) where CUDA is unavailable. Feature support tracks the driver; occasionally slower than a native path. |
| rocm | AMD GPU (Linux) | high | AMD's open compute stack mapped as an executor. Enables GPU execution on supported AMD cards under Linux. ROCm support is tightly tied to OS/kernel and GPU generation; verify against your release. |
| coreml | Apple Silicon / macOS | mid | Apple's on-device ML framework. Lets modern Macs run the heavy models on the GPU/ANE. Requires macOS + CoreML capable hardware; model coverage can lag other providers. |
| openvino | Intel CPU / iGPU | low | Intel's optimized runtime. A solid CPU/iGPU path when no discrete accelerator is available. Good default for Intel laptops without a supported dGPU. |
| cpu | Any | low | Runs every model on the host processor. Always available as a bottom line; slowest for the heavy networks. Tune --execution-thread-count to your core count to avoid thrashing. |
You can pass more than one provider - Facefusion falls through to the next that your hardware supports. Exact supported set and order are version- and install-dependent. Earlier releases used --execution-provider (singular); current ones use --execution-providers.
Threads & queues
Use the provider list for raw execution; use --execution-thread-count to cap per-request worker threads and --execution-queue-count to bound queued work between the video/audio pipelines. These are per-process knobs, not per-model accelerators - they gate CPU/queue pressure, not GPU clocks.
--execution-thread-count1 - 128 (version-dependent)
Kept low-ish by default and auto-tuned in some releases; setting it above your physical cores rarely helps and can hurt. Watch --help for the current clamp.
--execution-queue-count1 - 32
Higher smooths frame pacing at the cost of RAM; keep small on tight systems.
VRAM budgeting
Facefusion lets you cap per-provider GPU memory so several models can share one card without spilling. In current releases the argument is --vram-limit (and --system-memory-limit for host RAM). A strict budget forces the runtime to subdivide model loads and run frames in smaller passes.
--vram-limit- Exact units, ranges and defaults drift between versions (and by provider), so treat any figure here as qualitative. Verify with --help.--system-memory-limit- caps host RAM used by the pipeline.--max-memory- Earlier releases called this --max-memory; the name changed.
A representative run
# NVIDIA: try the fast paths first
ff-run \
--execution-providers cuda \
--execution-thread-count 6 \
--execution-queue-count 2 \
--vram-limit 8 \
--processors face_swapper,face_enhancer \
--source <source> --target <target> --output <out>
# Laptop fallback: CoreML on Apple, DirectML on Windows, OpenVINO/cpu on Intel
ff-run \
--execution-providers coreml,directml,cpu \
--execution-thread-count 4 \
--processors face_swapper \
--source <source> --target <target> --output <out>Provider ordering, --vram-limit syntax, and whether the queue defaults differ between cores are all version-dependent. Read your build's ff-run --help for exact flags.