FacefusionField Guide

Encoding & pipeline

Render it right, and fast

The last stage is where results become files. Choosing a software or hardware encoder, setting a CRF target, and deciding how to chain heavy multi-stage pipelines makes the difference between a quick preview and a painful one.

Hardware vs. software encoders

Hardware: h264_nvenc, hevc_nvenc, h264_amf, hevc_amf, h264_qsv, hevc_qsv, h264_v4l2m2m. Software: libsvtav1, libvpx-vp9, libx264, libx265.

EncoderTypeTrade-off
h264_nvenc / hevc_nvencHardware · NVIDIAOffloads encode to the GPU - the fastest path on NVIDIA cards. Larger files at the same CRF as software.
h264_amf / hevc_amfHardware · AMDAMD's GPU encoder; good speed on AMD GPUs.
h264_qsv / hevc_qsvHardware · IntelIntel Quick Sync; uses the iGPU's media engine.
libx264 / libx265SoftwareEncoder runs on CPU. Slower to encode but better compression - the guide's pick when quality-per-MB matters. libx265 (HEVC) compresses better than libx264 at the cost of more encode time and less universal playback.
libsvtav1 / libvpx-vp9SoftwareModern codecs - excellent compression, the most encode time. Good for archival distribution.

--output-video-codec is the current flag name. --output-video-encoder is the older name - one or the other depending on your release.

Quality - think CRF, not bitrate

--output-video-quality is the constant-rate-factor: a perceptual quality target, not a flat bitrate. Lower numbers are better.

18 - 20

Recommended

The guide's sweet spot: visually transparent on most footage without bloated files.

low (toward 0)

Higher quality

Closer to lossless; big files, slow encodes.

high (toward 51)

Smaller files

Visible artifacts; keep for drafts only.

Prefer --output-video-quality over --output-video-bitrate unless you must hit an exact file size. x264 presets (ultrafast..veryslow; default ~veryfast). Slower presets = better compression per bit. Hardware encoders expose their own preset scales.

Post-processing: grain & intermediates

  • --output-video-grain - Adds subtle film grain to mask banding / make output match source footage, especially after strong upscaling. Especially valuable after frame_enhancer smooths a frame.
  • --temp-frame-format - Intermediate frame format (e.g. png vs jpg). PNG = slower but lossless intermediates; jpg = smaller temp disk.
  • --keep-temp - Keep temp frames on disk for frame-level inspection instead of deleting them.

Multi-stage chaining

For a one-pass piped pipeline, comma-list the processors - Facefusion runs them in order per frame. For heavy multi-stage work (swap, then enhance, then lip-sync/expression restore), split into separate runs so each stage has its own tuned options and you keep lossless intermediates.

single-pass pipelinebash
ff-run \
  --processors face_swapper,face_enhancer,lip_syncer \
  --face-swapper-model inswapper_128 \
  --face-enhancer-model codeformer \
  --face-enhancer-blend 0.8 --face-enhancer-weight 0.5 \
  --lip-syncer-model wav2lip \
  --audio <audio> \
  --source <source> --target <target> --output <out>
two-stage for max qualitybash
# stage 1: swap only, keep lossless temp frames
ff-run --processors face_swapper \
  --face-swapper-model inswapper_128 \
  --temp-frame-format png --keep-temp \
  --source <source> --target <target> --output <stage1>

# stage 2: enhance the swapped frames into the final encode
ff-run --processors face_enhancer,frame_enhancer \
  --face-enhancer-model codeformer --face-enhancer-blend 0.85 \
  --frame-enhancer-model real_esrgan_x4plus \
  --output-video-codec libx265 --output-video-quality 18 \
  --source <source> --target <stage1> --output <final>

The --audio flag and multi-run inputs are shown as examples; exact chaining options and whether a stage accepts frames or a video as its target depend on your release. Keep the --source identity handy for every stage that needs it.

cli.py