Transcribe Audio Files to Text on Mac: On-Device, No Upload
Turn MP3, M4A, WAV, or AIFF files into accurate text on your Mac — fully on-device, no upload, no per-minute fees. The fastest private workflow in 2026.
Dictanta imports and transcribes existing audio files on your Mac — no bot in the call, free for your first 3 recordings, then from $9.99/mo.
Download on the App Store Dictanta for Mac, iPhone & iPadYou have the file already. An interview saved as an MP3, a lecture someone recorded on their phone, a two-hour M4A of a client call, a Voice Memo from a walk where you talked through an idea. The audio exists; what you need is the text. And when you search for how to transcribe audio files to text on a Mac, the top results mostly want the same thing from you: upload the file to our server, make an account, and pay by the minute.
For a lot of files that’s a bad trade. Uploading a two-hour recording on hotel Wi-Fi is slow. Paying per minute punishes exactly the long recordings that most need transcription. And for anything sensitive — a client call, a medical appointment, an interview a source gave you on background — “send it to our cloud” is the wrong answer before price even comes up.
The good news is that in 2026 the Mac itself is a fast, accurate transcription machine. This post covers the three realistic ways to get text out of an audio file on a Mac, and walks through the on-device path in detail.
The three ways to transcribe an audio file on a Mac
1. Cloud upload services. Otter, Rev, Descript, and dozens of browser tools accept uploaded files. Accuracy is good, and human-reviewed services like Rev’s premium tier are still the ceiling for messy audio. The costs: upload time, per-minute or subscription pricing, an account, and your audio sitting on someone else’s infrastructure. Otter’s free tier, for instance, caps total monthly minutes and per-conversation length, so a single long recording can blow through it.
2. Local Whisper tools. Apps like MacWhisper run OpenAI’s open-source Whisper model on your own machine. Genuinely local transcription, well-liked, and a fair choice — we’ve written a full comparison. Two caveats: the polished AI-summary features typically call a cloud API by default (local models are possible with extra setup), and Whisper-family models are no longer the speed benchmark on Apple silicon.
3. Apple’s on-device speech stack. macOS 26 ships SpeechAnalyzer, Apple’s speech-recognition framework that runs entirely on the Mac. In MacStories’ independent testing (June 2025), it transcribed roughly 55% faster than Whisper Large V3 Turbo on the same hardware. Apps built on it — Dictanta is one — transcribe an imported file on-device with no account and no upload, and pair the transcript with an on-device summary.
The rest of this post is the third path, because for the common case — a clean-enough recording, a privacy or cost reason not to upload — it’s the shortest distance between “I have a file” and “I have the text.”
Transcribing a file on-device, step by step
Assumes a Mac running macOS 26.4 on Apple silicon (M1 or later), with Dictanta installed from the Mac App Store.
1. Drag the file in (or press ⌘O)
Drop the audio file anywhere onto the Dictanta window, or hit ⌘O and pick it from the
file dialog. M4A, MP3, WAV, and AIFF all import directly, along with other common
formats — that covers Voice Memos, podcast downloads, most dictation and recorder
apps, and files exported from editing software. Voice Memos recordings in particular
import as-is; there’s a dedicated walkthrough
if that’s your source.
One detail worth knowing before you commit to anything: imports don’t consume the free quota. Dictanta’s free tier covers your first 3 completed recordings, but importing an existing file isn’t a recording — so you can run your actual backlog through it and judge the transcript quality on your own audio before any pricing question comes up.
2. Let the batch pass run
Live recordings in Dictanta transcribe as they happen; imported files instead get a batch transcription pass that starts on import. The work happens on the Mac’s own silicon. If the language model for your audio’s language isn’t on the Mac yet, macOS fetches it once, on demand — after that, transcription runs fully offline. You can verify this the blunt way: turn off Wi-Fi and import another file. It transcribes anyway.
Nothing about the file is sent anywhere. There’s no server-side queue, no “processing in the cloud” spinner, no vendor with a copy. For a recording you promised someone would stay private, that’s the entire argument.
3. Read the transcript — and the summary
When the pass finishes you get two artifacts, both generated on-device:
- A timestamped transcript. Every segment is linked to its spot in the audio — click a line and playback jumps to that moment. This matters more for imported files than for anything else, because with a file someone else recorded, you weren’t in the room: checking what was actually said around an ambiguous line is a click, not a scrub through a two-hour waveform.
- A structured summary, generated by Apple’s on-device Foundation Models: a TL;DR, action items with best-effort owner and due-date inference, decisions, and open questions. Each summary bullet is audio-anchored — click it and hear the span it came from. For a lecture or interview, the TL;DR alone tells you whether the file deserves a full read.
The recording is titled with a timestamp at first, then auto-renames itself from the summary topic — so “New Recording 47.m4a” becomes something you can actually find later. Speaking of finding: Dictanta’s search covers titles, transcripts, and summaries, in-app and via Spotlight. Import a semester of lectures and “the one where she explained backpropagation” becomes a search box away.
4. Correct, then export
No transcription — cloud, Whisper, or Apple’s — is perfect, and imported audio is often rougher than a well-miked meeting. The transcript is editable in place, and the audio anchoring makes corrections fast: click the doubtful line, listen, fix.
Then get the text where it needs to go. Export formats are Markdown, PDF, Word (.docx), SRT, and JSON:
- Markdown pastes cleanly into Notion, Obsidian, Bear, or Apple Notes.
- Word is what editors, lawyers, and clients expect back.
- SRT turns a podcast or video’s audio into subtitles — there’s a separate guide to the SRT workflow if captions are the goal.
- JSON carries the timestamps for anything programmatic.
If the recipient just needs the gist, “Share Recap” produces a clean summary message without the full transcript attached.
Where the on-device path is strongest
Long files. Per-minute cloud pricing makes a three-hour recording expensive to even try. On-device, length costs you nothing but local processing time — and there’s no per-minute meter on any Dictanta tier.
Batches. Ten interview files for a research project means ten uploads on the cloud path. Locally it’s one multi-select drag, and the results are searchable together.
Sensitive audio. Interviews on background, client calls, anything covered by an NDA or a promise. The honest answer to “where did that file go” is “nowhere” — which is the same reason the on-device stack wins for live meeting capture.
No-network situations. On a flight or a locked-down corporate network, the offline path isn’t degraded — it’s identical, because the network was never part of the pipeline.
Foreign-language reading. If the file is in one of eight supported languages (Spanish, French, German, Italian, Portuguese, Japanese, Korean, Chinese Simplified), Dictanta can show an on-device translated reading view of the transcript. It’s ephemeral — a reading aid, not a persisted or exportable document — but for “what is this Spanish-language interview about,” it answers the question without a copy-paste trip through a translation website.
Not just the Mac
The same import path works on iPad — drag a file in from Split View or the Files app and the batch pass runs on the iPad’s own silicon. And whichever device did the transcribing, the results don’t stay siloed: transcripts, summaries, titles, and your edits sync across your Apple devices through your own iCloud, so a lecture transcribed on the iPad is searchable from the Mac an hour later. The audio itself never syncs — it stays on the device that holds the file, which also means audio-anchored playback works on that device.
Honest limits
- No speaker labels. The transcript doesn’t tag who said what; diarization is planned but not shipped. For a two-person interview you can usually follow the turns; for a six-person panel, a service with diarization may earn its upload.
- Audio quality still rules. A phone in a pocket at a noisy dinner defeats every transcription engine. On-device transcription is competitive with cloud engines on reasonable audio, but it can’t recover words that were never captured.
- Very rough audio may justify humans. If the recording is bad and the stakes are high — legal, broadcast — a human-reviewed service is still the accuracy ceiling, at human prices.
- Language coverage is finite. Apple’s speech models cover the major languages and the OS fetches them on demand, but if your audio is in a language the system doesn’t support, this path ends there.
When a cloud service is still the right call
Being fair to the alternatives: if you need per-speaker attribution today, if the audio is rough enough to need human review, or if your workflow lives inside a specific cloud editor like Descript’s, then uploading is the price of those features and it may be worth paying. The on-device path is for the much more common case — decent audio, a text destination like Notion or Word, and no good reason to hand a copy of the recording to a third vendor.
Bottom line
To transcribe audio files to text on a Mac in 2026, you no longer have to pick between uploading to a per-minute cloud service and wrangling command-line Whisper builds. Apple’s on-device speech stack handles M4A, MP3, WAV, and AIFF files locally, fast, and offline once the language model is fetched. Dictanta wraps that stack in a drag-and-drop import, adds an on-device summary with audio-anchored bullets, and exports to Markdown, PDF, Word, SRT, and JSON — and because imported files don’t touch the free quota, you can test it against your real backlog before spending anything. The files are already on your Mac. The transcription might as well happen there too.
FAQ
How do I transcribe an audio file to text on a Mac?
Drag the file into an on-device transcription app like Dictanta, or press ⌘O and pick it. The Mac transcribes the audio locally using Apple’s speech recognition — no upload, no account, no per-minute charge. You get a timestamped transcript you can search, correct, and export as Markdown, PDF, Word, SRT, or JSON.
Can I transcribe audio files on a Mac without uploading them anywhere?
Yes. Apple ships an on-device speech framework called SpeechAnalyzer in macOS 26, and apps built on it transcribe entirely on the Mac’s own Apple silicon. The OS downloads a language model once if it’s missing; after that, transcription works with the network cable unplugged. The audio file never leaves the machine.
What audio formats can be transcribed on a Mac?
Common formats work directly: M4A, MP3, WAV, and AIFF, among others — which covers Voice Memos exports, podcast downloads, dictation apps, and most recorder output. If a file is in something exotic, converting it to M4A or WAV first with a free tool solves it, but for typical recordings no conversion step is needed.