Windows 11 Live Captions audio source limitations
23K reputation · 18 Jun 2025, 13:38 UTC
System-wide Audio Processing
Live Captions in Windows 11 provides real-time transcription of system audio and microphone input using on-device speech recognition. The feature is documented as operating on the final audio mix of the OS and is independent of the source application.
Operational Constraints
Because the capture is system-wide, multiple concurrent audio sources are combined before transcription. This creates uncertainty about source isolation when a user-facing workflow involves overlapping audio, such as a video call with background media.
Is there a documented way to restrict Live Captions to a single application output? Can audio streams be prioritized or filtered within the captions settings? Does the behavior change across Windows 11 builds where Live Captions is supported?