Private captioning guide

Local AI captions for screen recordings on Mac

Captions make screen recordings easier to follow in silence, more accessible and more resilient to imperfect audio. Running transcription locally adds another benefit: the voice track does not need to leave your Mac.

In short: Local transcription processes the audio on your Mac instead of uploading it to a captioning service. Cadre uses Whisper through whisper.cpp to create word-timed segments, then renders styled captions into the exported video.

Why local captions matter

Many product videos contain internal roadmaps, customer-shaped examples, unreleased interfaces or engineering details. Uploading the audio to a third-party transcription service creates another copy and another policy to evaluate.

On-device transcription keeps the audio processing on the Mac. It also works without waiting for an upload and makes the captioning workflow available even when the recording should stay offline.

How on-device transcription works in Cadre

  1. Cadre extracts and mixes the relevant audio from the project.
  2. A local Whisper model transcribes it through whisper.cpp.
  3. The result returns as timed segments with word-level timing.
  4. You review the words, line breaks, timing and visual style.
  5. The captions are rendered into the MP4 together with cuts, zooms and mixed audio.

The model runs on the Mac. The captioning step does not send the recording's audio to Cadre or a hosted transcription API.

Edit captions for reading, not transcription

A transcript records what was said. A caption guides a viewer through when it was said. Good captions need editorial decisions:

  • Break lines at natural phrase boundaries.
  • Keep each caption on screen long enough to read once comfortably.
  • Remove filler only when doing so does not change the speaker's meaning or character.
  • Use punctuation to clarify pace, not to imitate every breath.
  • Keep captions clear of menus, buttons and the area currently being demonstrated.
  • Use strong contrast and a size that remains legible on a phone.

Improve caption accuracy before you press record

  • Use a consistent microphone distance and reduce room echo.
  • Wear headphones so system audio does not leak into the narration track.
  • Speak product names, acronyms and technical terms cleanly.
  • Pause briefly between sections rather than speaking over loading states.
  • Review names, numbers, URLs and negations carefully; those small errors change meaning.

No transcription model removes the need for review. The useful automation is the first pass: timing every word so your attention can go to meaning and readability.

Know exactly what stays local

“AI-powered” does not automatically mean cloud-based. In Cadre, the recordings, project manifest and caption transcription stay on the Mac. Agent editing also connects to Cadre over a local loopback server. The agent itself may have its own data policy, so evaluate the agent separately and avoid pasting sensitive material into prompts.

For the full product-level explanation, read Cadre's privacy page and the guide to private screen recording.

Turn the recording into the finished story

Cadre records your Mac and gives you an editable cinematic timeline—automatic zooms, local captions, cuts, styling and an AI agent that can follow your direction.

Record, edit and preview free · Cadre Pro required to export · Apple Silicon · macOS 13+

Frequently asked questions

Do Cadre captions upload my audio?

No. Cadre runs Whisper locally through whisper.cpp. The captioning process extracts and transcribes audio on your Mac.

Are local AI captions accurate?

They can provide a strong timed first pass, but every transcription should be reviewed. Names, numbers, acronyms and words spoken over noise deserve special attention.

Can I style the captions?

Yes. Cadre stores captions as editable project data and renders the chosen caption styling into the finished MP4.

Why use burned-in captions?

Burned-in captions remain visible wherever the video is posted, including feeds and players that do not preserve a separate caption track. They are less flexible than an optional subtitle track, so keep a clean transcript when you also need platform-native captions.