How to Transcribe Zoom Meetings on Mac — No Bot, No Cloud
Transcribe Zoom calls on your Mac with no bot in the participant list and no cloud upload — on-device with Apple's speech models. Free for your first 3 meetings.
Dictanta captures the call's audio straight from the menu bar on your Mac — no bot in the call, free for your first 3 recordings, then from $9.99/mo.
Download on the App Store Dictanta for Mac, iPhone & iPadYou are in three Zoom calls today, two tomorrow, and at some point this week one of them is the one you actually wanted notes for — the pricing review, the customer interview, the board prep. You opened Otter once. Fireflies sent you a calendar invite from a bot named “Fred.” Granola skips the bot but still ships your audio to the cloud for processing. None of those are wrong. They are just very visible — or very cloud.
There is a reason a lot of people would rather not have an attendee in their meeting whose only job is to listen, transcribe, and ship the audio off to a third-party server. Some of those reasons are about privacy. Some are about the awkwardness of explaining to a customer why their words are being recorded by a Zoom participant named “Otter.ai Notetaker.” Some are about a compliance review.
This post walks through the alternative: how to transcribe a Zoom meeting on your Mac without any bot in the call, with the audio never leaving your laptop, using Apple Intelligence and Dictanta. If you have used Otter or Granola before and want to know what changes, the short answer is: nothing visible to the other side of the call.
What “without a bot” actually means
When you connect Otter, Fireflies, Fellow, Read.ai, or any of the cloud-bot transcription services to your Zoom calendar, what they do is roughly this:
- Watch your calendar for a meeting URL.
- Spin up a participant — a virtual attendee — and have it join the Zoom call as a user.
- That participant streams the meeting audio to the vendor’s servers in real time.
- A cloud ASR system (Whisper, AssemblyAI, Deepgram, or proprietary) transcribes it.
- A cloud LLM (GPT-4o, Claude, Gemini) summarizes the transcript.
- The transcript and summary land in your dashboard.
That works. It is also exactly six places your meeting audio gets handled by software other than Zoom itself. And every other participant in the meeting sees the bot show up in the participant list. Some hosts disable that for executive calls. Some legal teams flat-out forbid it. Some customers ask, on the call, “who’s Fred?”
The alternative is to capture the audio your Mac is already hearing — the audio that comes out of Zoom and into your speakers — and transcribe it locally. No bot. No second participant. No data leaves the machine.
On macOS 26 (Tahoe), Apple shipped two APIs that finally make this trivial:
- ScreenCaptureKit with audio-only capture. Lets a sandboxed app subscribe to the system audio output of a specific running process — like Zoom, or Teams, or Google Meet in Safari.
- SpeechAnalyzer, the new on-device automatic speech recognition framework introduced at WWDC 2025. In MacStories’ independent testing it transcribed English audio about 55% faster than Whisper Large V3 Turbo on the same Apple silicon. Critically, it runs entirely on-device. No network.
Combined, they let an app like Dictanta do the bot’s job — without being a bot.
The end-to-end flow
Here is what actually happens when you record a Zoom meeting with Dictanta on a Mac running macOS 26.
1. Start the meeting normally
Open Zoom. Join the call. Don’t enable Zoom’s built-in cloud recording (you can if you want; they are independent). Don’t invite any third-party participant.
2. Click Record in the menu bar
Dictanta lives in your Mac menu bar. Click the waveform icon and hit Record. The icon turns
coral and pulses, and the menu-bar panel shows a live caption of what is being said in real
time. When Dictanta or its panel is frontmost, ⇧⌘R starts a new recording from the keyboard;
if you want to start or stop recording while Zoom has focus, enable the optional global record
shortcut (⌃⌥⌘R) with the “Global record shortcut” toggle in Dictanta’s Settings — it’s off
by default.
The first time you record system audio, macOS prompts you for screen recording permission (audio capture is gated by the same TCC permission as screen recording — that’s the API contract, even though no video is captured). You grant it once.
3. Dictanta auto-detects the meeting app
Dictanta’s source picker has exactly three options: My voice (just your mic), This meeting (the call audio of the meeting app Dictanta detected, plus your mic), and Another app (the frontmost app’s audio, plus your mic). If Zoom is running, “This meeting” already points at it — same for Teams, Google Meet, and Webex. One click and the capture is scoped to the call; “Another app” covers everything else that plays audio, with your spoken commentary mixed in.
Recordings start out titled with a timestamp, then rename themselves after the meeting topic once the summary lands; the library row shows the date, duration, and source type.
4. Live partial captions stream during the call
Two things are useful in-meeting, not just after:
- A live caption line in the menu-bar panel, streaming what’s being said.
- A live waveform in the same panel, so a glance confirms the capture is running.
These are not transcription previews to be edited later. They are SpeechAnalyzer’s real-time partial hypotheses, streaming in near real time as people speak. (The full transcript view opens after the call — double-click a recording in the library, or right-click it and choose Open in New Window.)
If the meeting is in a language other than English, Dictanta asks SpeechAnalyzer for the model matching that locale and falls back to English if your Mac has no match. English is the most mature, and coverage grows as Apple ships new language models with macOS.
5. Stop recording, get a transcript and summary
When you end the meeting (or click “stop” in Dictanta), three things happen, all on-device:
- The full transcript is finalized with timestamps. Each segment is tap-to-seek.
- Apple’s Foundation Models LLM generates a summary: TL;DR, decisions, action items (with owners and rough due-date guesses), and open questions.
- Every summary bullet is anchored to the audio span it came from. Hovering a bullet highlights the corresponding waveform segment; clicking scrubs the audio to that moment and highlights the transcript line.
The last point is the one that matters most. If you don’t trust an AI summary, you can verify any bullet in one click. The audio is the source of truth — the LLM’s job is just to organize it.
6. Export, or just keep it on the Mac
By default the transcript and summary sync to your other Apple devices via CloudKit — transcripts and summaries only. The audio never syncs; there is no setting that sends it anywhere.
For export, the options are:
- Markdown — clean structure, drops straight into Notion, Obsidian, Apple Notes, Bear, or any markdown editor.
- PDF — for sharing a formatted transcript with someone who just needs to read it.
- Word (.docx) — for teams that live in Word or Google Docs.
- SRT subtitles — timestamped caption blocks for video workflows.
- JSON — a lossless dump of every field, for tooling integrations.
The audio file itself stays on your Mac until you delete the recording — delete a recording and its audio is gone with it. The transcript and summary stay until you delete them.
What this lets you skip
If you have been using a cloud-bot transcription service, here is the short list of things that go away:
- The visible bot in the participant list. Customers don’t ask who Fred is anymore.
- The “we’re going to record this call for note-taking” disclosure — you can still say it for ethical reasons, but it isn’t required by any third-party ToS because no third party is involved.
- The vendor’s data-retention policy. Your call audio is on your Mac, scoped to your user account, encrypted by FileVault (if you have it on, which you should). When you delete it, it’s deleted.
- The pricing meter. Otter’s free plan caps at 300 minutes per month and 30 minutes per conversation. Granola’s free plan has unlimited meetings but a limited note-history window; paid starts at $14/user/mo. Fireflies Pro is $10/seat/mo billed annually ($18 month-to-month), with AI-credit limits on top. (Prices checked September 2026.) Dictanta is free for your first three meetings, then $9.99/mo, $79.99/yr, or $149.99 lifetime — no per-minute meter at any tier.
What about Zoom’s own transcription?
Zoom ships its own answer to this, and for some accounts it’s genuinely fine: cloud recording produces a transcript, and AI Companion generates a meeting summary at no extra cost — on paid plans. The dependency list is the catch. AI Companion isn’t available on Zoom’s free Basic tier, it’s controlled at the account level (your admin has to enable it, and many lock it down), and both the recording and the summary are processed and stored in Zoom’s cloud. If you’re a guest in someone else’s meeting, their account settings decide whether you get anything at all.
The on-device path has none of those dependencies: it works on a free Zoom account, in meetings you don’t host, with nothing enabled on Zoom’s side, because the capture happens entirely on your Mac. The full comparison is in the Zoom AI Companion alternative write-up — the short version is that AI Companion is worth using when your admin has it on and cloud processing is acceptable, and worth routing around when either of those is false.
Where this approach has limits
To be straight: there are a few things the bot approach does that the no-bot approach does not.
- Bots tag speakers automatically because they see who is speaking in Zoom’s participant list. Dictanta does not ship speaker diarization (separating who said what) — Apple’s on-device speaker-ID for system audio is not yet production-grade on short captures. Diarization is planned for a later release.
- Bots can join meetings you are not in. If your assistant is in a meeting you skipped, Otter can still capture it. Dictanta requires you to be at the Mac running the recording.
- Bots can integrate with Slack, HubSpot, Salesforce out of the box. Dictanta exports to Markdown, PDF, Word, SRT, and JSON, and offers no built-in CRM, calendar, or webhook integrations. If “the CRM updates itself after every call” is the requirement, the cloud-bot category is where that lives.
If those are important to you, the cloud-bot category is the right answer. If you have ever been the executive on a customer call wishing the third-party note-taker wasn’t there — or if your legal team has ever asked where the recordings live — the no-bot path is the right answer.
A note on macOS permissions
The first time Dictanta records, you grant it screen recording permission. The system shows a sheet that says it lets the app capture your screen and system audio. Dictanta does not capture your screen — it asks for the permission because Apple’s ScreenCaptureKit API gates audio capture behind the same TCC entitlement as screen capture, even when no video frames are requested. (This is a known Apple API design choice; same for every Mac transcription app that uses ScreenCaptureKit.)
The capture is audio-only by construction: the ScreenCaptureKit content filter includes
only the meeting apps you choose, and Dictanta attaches only audio (and optionally
microphone) stream outputs to it — never a video output — so no video frames are ever
consumed. macOS offers no audio-only permission slot for this API, which is why every
ScreenCaptureKit-based Mac transcription app hits the same screen-recording gate. If you
want to check for yourself, the installed app’s own Info.plist
(Dictanta.app/Contents/Info.plist) carries capture usage strings stating that Dictanta
records only audio and that it stays on the Mac.
If your IT department has a managed Mac and disables screen recording entitlements, Dictanta falls back to mic-only recording (so you can still record meetings, just only what your mic hears — useful for an in-person meeting, less useful for remote calls).
Trying it in your next meeting
Practical workflow for someone trying this for the first time on a real call:
- Install Dictanta from the Mac App Store on your work Mac. The download is small (~4 MB plus the SpeechAnalyzer language model, which macOS fetches once if it isn’t already on the Mac; transcription then runs fully offline).
- Open Dictanta once. Grant microphone and screen recording permissions when prompted.
If you want to start recordings while Zoom has focus, flip on the “Global record
shortcut” toggle (
⌃⌥⌘R) in Settings. - In your next Zoom call, click Record in the menu bar (or press the global shortcut, if you enabled it) when the meeting starts. Forget about it.
- After the call, open Dictanta. The recording will be on the meetings list. Tap into it to see the transcript and summary.
- If you want it in Notion or Obsidian, click Export → Markdown.
That’s the whole flow. No vendor onboarding, no calendar OAuth, no bot tested in a sample call. The free tier covers three meetings, which is enough to see whether the summary quality is good enough for what you need. If it is, upgrade; if not, you spent zero money.
Bottom line
For Mac users who do a lot of Zoom (or Teams, or Google Meet, or Webex) calls and don’t want to send their meeting audio through someone else’s pipeline, the right tool is the one that captures the audio your Mac is already hearing, transcribes it on-device with Apple’s SpeechAnalyzer, and summarizes it on-device with Apple’s Foundation Models. No bot in the call, no cloud upload, no monthly meter.
That tool is Dictanta. It requires macOS 26.4, iOS 26.4, iPadOS 26.4, or visionOS 26. The Mac version is the most feature-complete because system-audio capture only exists on the Mac.
If you’re shopping for a transcription tool right now, the question isn’t whether the cloud options are good — they are. The question is whether you want the cloud at all for this particular workflow. If the answer is no, you finally have an option that respects that.
FAQ
Can I transcribe a Zoom meeting on a Mac without a bot joining the call?
Yes. On macOS 26, Dictanta captures the system audio your Mac is already hearing — the Zoom call coming out of your speakers — plus your microphone, and transcribes it locally with Apple’s SpeechAnalyzer. No second participant joins the meeting and no audio leaves the machine.
Does the other side of the call know I’m recording?
Nothing is visible to the other side. There is no bot in the participant list and no third-party service involved, so no vendor terms require a disclosure. Recording-consent law still applies, though: some jurisdictions require every party’s consent, so check your local rules — and announcing the recording is good practice regardless.
Does the meeting audio ever leave my Mac?
No. Capture, transcription, and summarization all run on-device. Transcripts and summaries sync to your other Apple devices via CloudKit; the audio never syncs — there is no setting that sends it anywhere. Delete a recording and its audio is gone with it.
What does Dictanta cost?
It’s free for your first three meetings. After that it’s $9.99/mo, $79.99/yr, or $149.99 lifetime — with no per-minute meter at any tier.