Transcribe FaceTime Calls on Mac: On-Device, No Cloud
FaceTime has no record button on the Mac and no saved transcript. How to capture and transcribe FaceTime calls locally — on-device, with consent, no cloud.
Dictanta captures the call's audio straight from the menu bar on your Mac — no bot in the call, free for your first 3 recordings, then from $9.99/mo.
Download on the App Store Dictanta for Mac, iPhone & iPadFaceTime is where some of the most consequential calls happen — the user-research interview with someone who doesn’t have Zoom, the freelance client who prefers it, the call with a parent’s doctor, the source who’ll only talk on a platform they trust. And FaceTime gives you nothing afterward: no recording, no transcript, no notes. If you’ve tried to transcribe FaceTime calls on a Mac, you’ve hit the wall everyone hits — the app has no record button, and none of the notetaker services support it.
The way through is the same local-capture path that works for Zoom, Teams, and Meet on a Mac: record the audio your Mac is already playing, transcribe it on-device, and keep every byte on the machine. FaceTime needs a couple of specific adjustments to that flow, which is what this post covers — along with the consent conversation you should have first, because FaceTime calls are personal in a way work calls usually aren’t.
What Apple gives you natively
Apple has been adding call intelligence, but the pieces don’t add up to a FaceTime transcript on the Mac:
- On iPhone, the Phone app can record calls. It announces the recording to everyone on the call and drops the audio and an on-device transcript into Notes. It covers calls the Phone app handles — regular phone calls and FaceTime audio placed through it. It does not exist on the Mac, and it doesn’t cover a FaceTime video call you take in the FaceTime app. The broader lay of the land is in the Apple Intelligence transcription overview.
- Live Captions can caption FaceTime in real time. It’s an Accessibility feature, English-focused, and the captions are an overlay — nothing is saved. When the call ends, the words are gone. Captions and transcripts get conflated constantly; one is assistance during the call, the other is a record after it.
- The FaceTime app itself saves nothing. No recording feature on any Apple platform, video or audio, host or guest. SharePlay, screen sharing, and FaceTime links didn’t change that.
So the native answer on a Mac is: watch the captions scroll by, remember what you can.
Why the notetaker services are a dead end here
Cloud notetakers — Otter, Fireflies, Read.ai — reach meetings by joining them as a bot participant through a public join link. Zoom, Meet, Teams, and Webex all have a guest mechanism a bot can walk through. FaceTime has nothing like it: no bot role, no dial-in, no API for a service to attend a call. Even the browser-extension notetakers that piggyback on Google Meet’s caption stream have no FaceTime equivalent to read. None of the notetaker vendors list FaceTime support, because the platform gives them no way in.
There’s an upside buried in that: FaceTime calls are end-to-end encrypted, and no third-party service sits in the middle of them. Recording locally at your end keeps that property intact — the call still never transits anyone’s cloud. Shipping the audio to a transcription vendor afterward would give away exactly the guarantee that made FaceTime the right choice for the call.
The local path on a Mac
Two macOS frameworks, the same pair behind every no-bot flow on this blog:
- ScreenCaptureKit with an audio-only content filter captures the system audio of one specific process — here, the FaceTime app. The capture is scoped: Music, notification sounds, and a YouTube tab stay out. It’s audio only — never video, never a pixel of the screen, even though macOS files the permission under Screen & System Audio Recording.
- SpeechAnalyzer, Apple’s on-device speech recognition framework, transcribes the stream live during the call. Fully offline once the language model is on the Mac.
Dictanta wires those together with Apple’s Foundation Models for the post-call summary. One FaceTime-specific detail: Dictanta auto-detects the major meeting apps — Zoom, Teams, Webex, Slack, and browsers — but FaceTime isn’t on the auto-detect list. That’s what the source picker’s third option is for. The picker has exactly three choices: My voice (just your mic), This meeting (an auto-detected meeting app’s audio plus your mic), and Another app (the frontmost app’s audio plus your mic). For FaceTime, you use Another app.
Recording a FaceTime call, step by step
Assume Dictanta is installed on macOS 26.4 or later on Apple silicon.
1. Ask first
Before anything else, because this is FaceTime: tell the other person you’d like to record and why. “I want to record this so I can transcribe it and not take notes while we talk — it stays on my Mac, nothing goes to any service” is a disclosure most people say yes to, and it’s more concrete than what any cloud notetaker can honestly offer. Recording-consent law varies by jurisdiction; many require all parties’ consent, and a personal call is exactly the context to hold yourself to the stricter standard regardless of what your local statute says.
2. Take the call on the Mac, and keep FaceTime frontmost
The call has to run on the Mac — if you answer on your iPhone, there’s no Mac audio to capture, and iOS doesn’t let apps capture other apps’ audio at all. If the call started on your iPhone, hand it off to the Mac before recording.
Because FaceTime isn’t auto-detected, Dictanta’s Another app option captures whichever app is frontmost when you start the recording. So: click the FaceTime window so it’s the active app, then start the capture. Which headphones or speakers you use doesn’t matter — the capture taps FaceTime’s audio stream directly, not the output device.
3. Start recording from the menu bar
Click Dictanta’s waveform icon in the menu bar, choose Another app as the source, and hit Record. Clicking a menu-bar panel doesn’t change which app is frontmost, so FaceTime stays the capture target. Your microphone is mixed in automatically — FaceTime’s audio is the other person, your mic is you, and the combined stream is the whole conversation. The panel’s live caption line and moving waveform confirm the capture is running.
The first system-audio recording triggers macOS’s one-time Screen & System Audio Recording permission prompt. Grant it once, never again.
4. Stop, and get the transcript and summary
When the call ends, stop the recording. On-device, three things land:
- A full transcript with per-segment timestamps — click a line, the audio jumps to that moment.
- A structured summary: TL;DR, action items with best-effort owner and due dates, decisions, and open questions. For a research interview or a call with a doctor, the open-questions section is often the most useful part — it’s the list of what you forgot to ask.
- Audio-anchored bullets. Every summary bullet links back to the audio span it came from, so “she said the dosage changes in September” is one click away from hearing exactly what was said. On a call where the details genuinely matter, checkable beats plausible.
The transcript autosaves every 30 seconds during the call, so a dropped call or a crash leaves a recovered recording, not nothing. The recording titles itself with a timestamp, then auto-renames from the summary topic.
5. Export it, or keep it in the library
Markdown, PDF, Word (.docx), SRT, and JSON export cover the downstream cases — Markdown into Obsidian or Notes for interview notes, Word for anything you need to send as a document. Or leave it in the library: on-device search covers titles, transcripts, and summaries, so “what did the contractor say about the permit” is findable months later. Transcripts and summaries sync to your other Apple devices through your own iCloud; the audio never syncs and never leaves the Mac.
FaceTime-specific things worth knowing
Audio and video calls capture identically. A FaceTime audio call and a FaceTime video call produce system audio the same way. Only the sound is ever captured.
Guests joining from a browser don’t change anything. FaceTime links let Windows and Android users join from a browser, but your end of the call still runs in the Mac’s FaceTime app, which is where the capture points. What the other side joins from is irrelevant.
Voice Isolation helps the transcript. FaceTime’s Voice Isolation mic mode strips background noise from what each side sends. Cleaner audio in means fewer transcription errors out — worth turning on for interviews.
Group FaceTime works, without speaker labels. A three-person call arrives as one mixed stream, so the transcript won’t tag who said what — same limitation as every system-audio path; speaker diarization is planned. For 1:1 calls this barely matters: your lines and theirs alternate, and your side came through your mic.
A call taken on iPhone or iPad leaves nothing to capture. System-audio capture is Mac-only. On those devices Dictanta records the mic with live captions — good for in-person conversations, not for pulling the far end of a FaceTime call.
Where this path has limits
- No speaker labels yet, as above.
- The frontmost-app rule is on you. Start the recording while some other app is frontmost and Another app captures that app instead. The live caption line gives it away immediately — if it isn’t transcribing the call, check the source and restart.
- You capture from when you press record. Nothing retroactive. For a scheduled interview that’s easy; for a call that turns important halfway through, you keep whatever you captured from the moment you noticed.
- It’s your record, not a shared one. Nothing posts anywhere automatically. For a work-interview workflow that’s usually the point; just don’t expect the other party to receive anything unless you send it.
Who actually needs this
The pattern shows up wherever FaceTime is the platform the other person trusts: user researchers interviewing participants who won’t install anything, journalists whose sources prefer an end-to-end-encrypted call, freelancers whose clients just like FaceTime, and anyone sitting in on a parent’s medical calls who needs the details right afterward. The common thread is that the call is sensitive enough that a private, no-cloud transcription path isn’t a nice-to-have — it’s the condition under which recording is acceptable at all.
Trying it on your next call
- Install Dictanta from the Mac App Store. macOS fetches the speech model once if needed; transcription then runs fully offline.
- Open it once and grant the microphone and system-audio permissions.
- On the next call: ask, click the FaceTime window, pick Another app, record.
The first 3 recordings are free with every feature unlocked and no length cap — enough to test it on a real call before paying anything. After that, Dictanta Pro is $9.99/month, $79.99/year, or $149.99 once for lifetime. Nothing is metered per minute; a two-hour interview transcribes the same as a five-minute check-in.
Bottom line
FaceTime gives you end-to-end encryption and zero record-keeping — the privacy is the point, and the amnesia is the price. Apple’s own tools don’t close the gap on the Mac: the Phone app’s recording lives on iPhone, and Live Captions saves nothing. Cloud notetakers can’t reach FaceTime at all. Local capture is the one path that closes the gap without giving up the privacy that made FaceTime the right choice: ScreenCaptureKit takes the call’s audio, SpeechAnalyzer transcribes it on-device, and the recording, transcript, and summary live on your Mac and nowhere else. Ask first, keep FaceTime frontmost, and the calls that matter stop evaporating.
FAQ
Can you record and transcribe a FaceTime call on a Mac?
Not with FaceTime itself — the Mac’s FaceTime app has no recording feature and no saved transcript. You can record the call locally instead: the Mac captures FaceTime’s audio output plus your microphone, and on-device speech recognition turns the combined stream into a transcript. Nothing is uploaded anywhere.
Does FaceTime notify the other person if you record the call this way?
No — the capture happens on your Mac, outside FaceTime, so FaceTime has nothing to announce. That makes disclosure your responsibility, and it matters more on FaceTime than on work platforms because the calls are often personal. Say you’re recording before you start; recording-consent law varies by jurisdiction and many places require everyone’s consent.
Why can’t Otter or Fireflies transcribe FaceTime calls?
Bot-based notetakers work by joining a meeting as an extra participant through a public join link, and FaceTime has no mechanism for that — there is no bot role and no API for one. Local capture on the Mac is the only practical way to get a FaceTime transcript.