Quick Answer
Realtime transcription is speech converted to text while you are still talking, streamed out word by word, instead of processed after the recording ends. The software slices incoming audio into fragments, prints a fast guess as an interim result, then revises it as more context arrives. Almost every realtime transcription product you will find is a hosted service: a streaming API a developer wires into an app, or a browser tab that holds an open connection to a vendor's servers. Both designs mean your audio leaves the machine while you speak. VoicePrivate is the on-device option for people rather than developers. It runs realtime transcription as live dictation on macOS 13 or later, types into whatever app has your cursor, reaches sub-200ms first-token latency on Apple Silicon, keeps working with Wi-Fi off after a one-time model download, and costs $84 a year or $9.99 a month with a 7-day free trial for individual monthly and annual plans and no free tier.
What Realtime Transcription Actually Means
Realtime transcription means the text appears while the speaking is still happening. That is the whole definition, and the rest follows from it.
Batch transcription waits. You hand it a finished file, it reads the entire thing with full context available, and it returns one settled transcript. Realtime transcription cannot wait for anything. It commits to a guess about the last half second of audio before it has heard the next half second, which means it prints words that are sometimes wrong and then quietly corrects them. If you have watched live captions on a video call flicker and rewrite themselves mid-sentence, you have watched that correction pass happen.
Two terms come up constantly and describe the same trade. Interim results are the fast, provisional words the engine emits immediately. Final results are what it settles on once it has heard enough to be confident, usually at a natural pause. A realtime system that shows only final results feels laggy but reads cleanly. One that shows interim results feels instant but visibly changes its mind.
There are two distinct jobs people mean by realtime transcription, and mixing them up leads to buying the wrong tool. One is live captioning: turning someone else's speech, a lecture, a call, a meeting, into text you read. The other is dictation: turning your own speech into text that lands in a document you are writing. VoicePrivate does the second job. It is dictation software for macOS, and the realtime part means the words show up in your active text field as you say them.
How Realtime Transcription Works, Slice by Slice
Every realtime transcription system, on-device or cloud, runs the same loop. The difference between products is where each step happens, not what the steps are.
- Capture. The microphone delivers a continuous audio stream, which the software chops into short chunks, typically fractions of a second.
- Feed the model. Each chunk goes into a speech model along with the recent context, so the model knows what was already said.
- Emit an interim guess. The model returns its best current reading of the words. This is what makes the output feel instant.
- Revise. As the next chunks arrive, earlier guesses get corrected. "recognise speech" turning into "wreck a nice beach" and back again is this stage doing its job badly.
- Detect the endpoint. The system looks for the pause that means you finished a thought, finalises that segment, and adds punctuation and capitalisation.
- Deliver. The finalised text goes wherever the product puts it: a caption overlay, a transcript pane, a websocket message back to a developer's app, or, in VoicePrivate's case, straight into the text field your cursor is in.
Step 2 is the fork in the road. In a hosted product, that chunk of audio is serialised and sent over the network to a vendor's servers, where the model lives. In an on-device product like VoicePrivate, the model is sitting on your own processor and the chunk never goes anywhere. Same loop, completely different answer to the question of who ends up holding a recording of the conversation.
Where Your Audio Goes During Realtime Transcription
Search realtime transcription and look at what actually comes back. The results are a developer API guide, a browser tool, cloud speech platforms, a roundup of hosted services, and a self-hosting thread on Reddit where people are asking for exactly the thing the rest of the page does not offer. Read that pattern honestly, because it tells you what the market is selling.
- Streaming APIs. OpenAI's realtime transcription guide walks developers through creating a transcription session, streaming audio into it, handling transcript events, and tuning latency against accuracy. Speechmatics markets its real-time product at 55 or more languages and advertises under a second of latency. Google Cloud Speech-to-Text lists streaming recognition, 85 or more languages and variants, and speaker diarization. These are well built and they are all the same architecture: an open socket carrying your audio to someone else's infrastructure.
- Browser tools. Maestra's live transcribe page offers realtime captions in a tab with a free mode and no sign-up. Convenient, and the audio is travelling to a server for the length of the session, free tier included.
- Roundups of hosted services. AssemblyAI's list of top live transcription tools names AssemblyAI, the OpenAI realtime API, Deepgram, Google Cloud Speech-to-Text, Speechmatics, Microsoft Azure AI Speech, Rev.ai, AWS Transcribe and Otter.ai. Every one of those is a hosted service. Not one of them processes the audio on your machine.
- Mobile accessibility apps. Google's Live Transcribe on Google Play captions the room around you on an Android screen. It is free, and it is a reading aid on a phone, not a way to get text into a document on a Mac.
None of those pages is dishonest. They are answering a developer's question or a captioning question. What none of them answers is the question a lawyer, clinician or advisor is actually asking: can I get words on screen as I speak without a recording of what I said crossing the network? VoicePrivate exists for that question, and the answer is architectural rather than contractual. There is no VoicePrivate server receiving your audio, so there is no retention policy to read and no vendor breach that could expose what you dictated.
Realtime Transcription That Never Leaves Your Mac
VoicePrivate does realtime transcription as live dictation on macOS 13 or later, on both Apple Silicon and Intel Macs. The speech model runs on your own processor. Here is what that buys you and what it costs you, concretely.
Text lands where you are already working
VoicePrivate types into whatever app has focus: a mail draft, Apple Notes, Slack, a code editor, a document in Pages, a form in a browser, a field in a practice management system. There is no separate transcript window to copy out of. You can also set per-app dictation modes, so VoicePrivate behaves one way in your email client and another in a clinical or drafting tool.
Latency with nothing in the path but your Mac
VoicePrivate reaches sub-200ms first-token latency in live dictation on Apple Silicon Macs. Because no network hop exists, that number is also stable: it does not degrade at 2pm when a vendor's servers are busy or when hotel Wi-Fi gets congested. Intel Macs are supported with the same accuracy, though first-token latency varies compared with Apple Silicon.
Offline after one download
VoicePrivate pulls its speech model down once on first launch. After that, realtime transcription works with the network off, permanently. No API key, no account, no session to re-establish on a plane or in a facility with no outbound internet.
Vocabulary you can fix yourself
Realtime transcription has no second pass, so a proper noun it mishears is a proper noun you retype. The VoicePrivate custom dictionary takes names, acronyms, drug names and case citations so the next session comes back cleaner. VoicePrivate covers this in detail in the guide to custom vocabulary for voice to text on Mac. VoicePrivate handles 99 languages, and for recordings that already exist on disk it also does file transcription with export to TXT, JSON, MD, SRT and VTT.
Flat pricing instead of a meter
Realtime speech APIs bill by streamed audio, which punishes whoever dictates most. VoicePrivate General edition is $84 a year or $9.99 a month with unlimited dictation, and a Team plan at $68 a year per seat for 2 to 5 floating seats. Individual monthly and annual plans include a 7-day free trial, a card is required and $0 is due today, and there is no free tier.
For the deeper latency breakdown on the dictation side, see VoicePrivate's page on real time voice to text on Mac, which covers first-token timing, macOS Live Captions, and how the built-in Apple Dictation compares.
Realtime Transcription Options Compared
These are the shapes of realtime transcription you will run into, sorted by the question that decides everything else: where does the audio get processed? VoicePrivate is listed first because it is our product and its row belongs next to everyone else's rather than on a page by itself.
| Option | Where audio is processed | Works with Wi-Fi off | Types into any app | Who it is for | Pricing model |
|---|---|---|---|---|---|
| VoicePrivate | On your Mac | Yes, after a one-time model download | Yes, at the cursor in any macOS app | Professionals dictating their own speech | Flat: $84 a year or $9.99 a month, 7-day free trial for individual monthly and annual plans |
| Realtime speech APIs | Vendor servers, over a streaming connection | No | No, they return data to your code | Developers building transcription into a product | Usage based, billed against streamed audio |
| Browser live transcribe tools | Vendor servers, for the whole session | No | No, text stays in the tab | Quick one-off captioning with nothing to install | Often a free mode, with paid tiers above it |
| macOS Live Captions (built in) | On your Mac on Apple silicon | Yes | No, it shows a floating caption window | Following audio you are listening to | Included with macOS |
| macOS Dictation (built in) | On your Mac for shorter passages on Apple Silicon | Yes, for on-device dictation | Yes, into the focused field | Short bursts of everyday typing | Included with macOS |
| Google Live Transcribe | Handled by the app on Android | Not verified here, check Google's listing | No, and there is no Mac version | Accessibility captioning on a phone | Free on Google Play |
| Meeting assistants such as Otter.ai | Vendor servers | No | No, transcripts live in the vendor workspace | Teams who want shared meeting transcripts | Subscription, usually per seat |
Rows describe product categories and the capabilities each vendor documents on its own site, checked September 2026. Competitor prices are described by billing model rather than by figure, because published rates change without notice.
The Latency Budget: What "Realtime" Has to Beat
Realtime transcription is not a binary. Everything has latency, and the only question is whether the delay is short enough that you stop noticing it. For dictation, the working threshold is roughly a quarter of a second. Past that, you start waiting for the software, and the moment you are waiting, you are thinking about the tool rather than the sentence.
Cloud realtime transcription carries a cost it cannot design away. The audio has to be packaged, sent across the network, queued, decoded and returned. Otter.ai documents a round trip of two to three seconds between the end of your speech and the first character appearing. That is fine for reading a meeting transcript afterwards and wrong for typing.
VoicePrivate removes every step in that chain except the decode. Live dictation reaches sub-200ms first-token latency on Apple Silicon Macs, which is inside the threshold where speech feels like typing. It is also predictable, because there is no shared infrastructure between you and the result. Nothing about your latency depends on how busy a vendor's cluster is.
One honest caveat about accuracy at speed. Every realtime system, including VoicePrivate, is less accurate than the same model given the whole recording, because a realtime system has to answer before it has heard the end of the sentence. If accuracy matters more than immediacy, record first and transcribe the file. VoicePrivate does both, and VoicePrivate's guide to audio transcription on Mac covers the file side.
Where Realtime Transcription Is the Wrong Tool
Realtime transcription is oversold, so here is where it, and VoicePrivate specifically, will let you down.
- A room full of people talking over each other. Three or more overlapping speakers is where every tool degrades, VoicePrivate included. VoicePrivate applies speaker labels for two speaker recordings. A busy conference table is not a solved problem for anyone.
- You need an API. VoicePrivate is a Mac application, not a streaming endpoint. If you are building realtime transcription into your own software, the hosted APIs in the table above are what that job needs.
- You need a shared live transcript. Several people watching the same transcript update in a browser is inherently a server-side feature. On-device processing and shared live documents pull in opposite directions, and VoicePrivate is firmly on the on-device side.
- You want captions of someone else's audio. That is macOS Live Captions under Accessibility, which shows a floating caption window. VoicePrivate is for dictating your own speech into your own documents.
- You need a certified transcript. Court reporting and other attested transcripts need a human who can stand behind the accuracy. No realtime engine replaces that.
- You are not on a Mac. VoicePrivate requires macOS 13 or later here. For Windows, see VoicePrivate's offline speech to text for Windows page.
Realtime Transcription in Regulated Work
The reason realtime transcription architecture gets attention in law, healthcare, finance and insurance is simple: with streaming, the audio leaves during the conversation, not afterwards when you had a chance to think about it. A cloud realtime session is an open pipe from your microphone to a third party for as long as you are talking.
VoicePrivate's answer is that no audio is transmitted at all. Processing happens on the device, so there is no upload to assess, no vendor-side copy of the dictation, and no realtime session log held anywhere but your own Mac. VoicePrivate ships editions carrying domain vocabulary on the same on-device architecture: Legal, Healthcare, Finance and Insurance. For the architectural comparison in full, VoicePrivate's breakdown of on-device versus cloud transcription goes through it step by step.
Realtime Transcription FAQ
What does realtime transcription mean?
Realtime transcription means speech is converted to text while the person is still talking, in a continuous stream, instead of after the recording is finished. The software takes audio in small slices, guesses at the words in each slice, prints that guess immediately as an interim result, and revises it a moment later once more context has arrived. Batch transcription does the opposite: it waits for the complete file and returns one settled transcript. Realtime buys you speed and costs you a second pass. VoicePrivate does realtime transcription as live dictation on a Mac, with the speech model running on your own processor rather than on a server.
Is Google Live Transcribe free?
Yes. Live Transcribe is a free Google accessibility app distributed through Google Play, and it captions the speech happening around you on the phone screen. Two limits matter if you found it while looking for realtime transcription for work. It is an Android app, so there is no Mac version, and it is a reading aid rather than an input method, so the text stays in its own window instead of landing in the document you are writing. VoicePrivate covers the other job on macOS: realtime transcription that types into whatever app has your cursor.
Can ChatGPT transcribe audio in real time?
Its voice features transcribe speech as you talk inside that app, and OpenAI publishes a realtime transcription API that developers can stream audio into. Both run in the cloud, which means the audio travels to a server before any text comes back, and neither one is a system-wide dictation tool that types into your mail client or your case management system. If your constraint is that a recording of the conversation must not leave your machine, a cloud assistant is the wrong shape of tool. VoicePrivate is a Mac app that does realtime transcription locally and inserts the text at your cursor in any application.
Will transcriptionists be replaced by AI?
The part that has clearly moved to software is the first draft. Realtime transcription is fast and cheap enough that paying a person to produce a rough transcript of clean audio is hard to justify. The parts that have not moved are the ones realtime transcription is worst at: crosstalk in a room full of people, heavy accents, poor recording conditions, speaker attribution that has to hold up to scrutiny, and certified transcripts where somebody has to attest to accuracy. The realistic shape is fewer keystrokes and more review work rather than a clean replacement. VoicePrivate is built for the first case, dictating your own speech, and is not a substitute for a certified human transcript.
Does realtime transcription work offline?
Only if the speech model runs on your own device. Streaming APIs and browser-based live transcription tools hold an open connection to a server for the whole session, so a dropped connection ends the session. VoicePrivate downloads its speech model once on first launch and then does realtime transcription with no network at all, which you can confirm in under a minute by switching Wi-Fi off and dictating a paragraph.
How fast does realtime transcription have to be to feel live?
Roughly a quarter of a second. Past that, you notice the wait and start pausing for the software instead of thinking about the sentence. Cloud tools carry a round trip that they cannot engineer away: Otter.ai documents two to three seconds between the end of your speech and the first character appearing. VoicePrivate reaches sub-200ms first-token latency in live dictation on Apple Silicon Macs because there is no network hop in the path, only the time the Mac needs to decode the audio.
What is the difference between realtime transcription and live captions?
They share an engine and differ in purpose. Live captions display someone else's speech on screen for you to read, which is what macOS Live Captions does in its floating window and what Google's Live Transcribe does on Android. Realtime dictation takes your speech and puts it into a document. VoicePrivate does the dictation job on macOS, inserting text at the cursor in any app rather than showing it in a window of its own.
How much does VoicePrivate realtime transcription cost?
VoicePrivate General edition is $84 a year or $9.99 a month, with a Team plan at $68 a year per seat for 2 to 5 floating seats. It is a flat fee with no per-minute meter, which is the main pricing difference from realtime speech APIs that bill by the hour of streamed audio. Individual monthly and annual plans include a 7-day free trial; a card is required and $0 is due today. There is no free tier. Specialty edition pricing is on the VoicePrivate pricing page.
Next Steps
- Dictation latency in detail: Real time voice to text on Mac covers first-token timing, Apple Dictation and macOS Live Captions.
- Recorded files rather than live speech: Audio transcription on Mac is VoicePrivate's guide to transcribing recordings that already exist.
- The architecture argument: On-device versus cloud transcription compares privacy, latency and accuracy directly.
- Coming from a cloud service: Cloud transcription software is an honest look at when cloud is the right call and when it is not.
- Teaching it your vocabulary: Custom vocabulary for voice to text on Mac covers names, jargon and acronyms.
- What the app does: VoicePrivate features and VoicePrivate pricing.
Bottom Line
Realtime transcription has been solved twice over, and almost always by shipping your voice to someone else's computer while you talk. If the content of what you dictate does not matter much, the hosted tools are good and you should use them. If it does matter, the only version of that promise you can verify yourself is the one where the model runs on your own machine, and the verification is turning off Wi-Fi and watching the words keep appearing. VoicePrivate is built for that case: realtime transcription on macOS 13 or later, typed into any app at sub-200ms first-token latency on Apple Silicon, $84 a year or $9.99 a month, no free tier.