Face detection is not speaker detection
Most auto-reframe tools detect faces and then pick one — usually the largest, the most centred, or the one moving most. That works on a single-presenter recording and falls apart the moment there are two people in frame, because the listener is often the more animated one. Nodding, reacting, and laughing all read as motion. Speaker detection asks a different question: not 'where is a face' but 'which of these faces is producing the audio right now'.
How FrameOS decides who is speaking
FrameOS combines the transcript's timing with per-moment audio and face tracking to build a timeline of who holds the floor across the recording. That timeline is what drives the framing decisions: when the speaker changes, the frame changes with it. Because the decision is made per moment rather than once for the whole clip, a three-minute exchange between two hosts is framed as an exchange rather than as a single fixed crop.
Why the cut timing matters as much as the choice
Getting the right person is only half the job. A frame that switches on every short interjection produces a clip that feels twitchy and hard to watch. FrameOS holds the frame through brief back-channel responses — the 'mm-hm's and short agreements that punctuate a real conversation — and cuts on genuine handoffs. The result reads like a camera operator made the call, not like a crop chasing a volume meter.
Where it applies
Speaker detection matters most on anything with more than one person on camera: two-host podcasts, remote interviews, panel recordings, recorded calls, and co-streamed content. On single-presenter footage it is largely redundant — there is only one candidate — which is why FrameOS applies it where multiple people are actually present rather than forcing it on every source.
What speaker detection does
- Builds a who-is-talking timeline across the whole recording.
- Drives framing from speech, not from whoever moves most.
- Holds through short interjections and cuts on real handoffs.
- Applies where multiple people are on camera; skipped where it adds nothing.
FAQ
What is speaker detection in video editing?
Identifying which person in a multi-person recording is talking at each moment, so downstream decisions — framing, cropping, captions — can follow the actual speaker rather than guessing from face position or motion.
Why does my auto-crop follow the wrong person?
Because most auto-crops pick a face by size, centrality, or motion. In a conversation the listener is frequently the most animated person on screen, so a motion-driven crop lands on them. Driving the choice from speech instead of motion is what fixes it.
Does speaker detection work on remote interviews?
Yes — remote interviews and recorded calls are a common case, including recordings where each person occupies a separate region of the frame.