Article

Live captions and transcripts as an accessibility tool

How live captions and recorded transcripts solve different problems for deaf and hard of hearing people, where each one helps, and what this kind of app is not.

Speech-to-text does two different jobs for deaf and hard of hearing people, and confusing them leads to disappointment. Live captioning puts words on a screen while someone is talking, so you can take part in the moment. A transcript is a record produced from a recording, so you can read afterwards what you could not follow at the time. Both are useful. Neither replaces the other, and neither replaces a human interpreter or a professional captioner when the stakes are high. This article explains where each one fits, what running the text on your own device changes, and, plainly, what a transcription app is not.

Live captions and transcripts are not the same thing

Live captions appear as the words are spoken. They are what you want in a live lecture, a conversation, a call, or a talk you are attending. Apple builds a system-level version into iPhone, iPad and Mac, and it works across apps and on audio playing on the device; a video conferencing tool usually offers its own captions inside a call; and professional CART, real-time captioning typed by a trained captioner, remains the accurate option where accuracy is not negotiable. Automatic live captions are fast and imperfect: they revise words as more context arrives, they struggle with crosstalk and heavy accents, and they rarely tell you who is speaking.

A transcript is produced from a recording. It arrives after the fact, and in exchange it is far more useful as a document: it can be read at your own pace, re-read, searched, quoted, corrected, exported, and shared with the people who were in the room. It can separate and name speakers. It can be summarized down to the decisions and the action items.

A practical setup for many people uses both: live captions to follow along in the moment, a recording running in parallel, and a transcript afterwards to fill in everything the live captions dropped. ThinkScribe is the second half of that pair. It records and transcribes on your own device; the transcript comes from the recording rather than streaming onto the screen word by word, which is exactly why it can be as thorough as it is.

Where an on-device transcript helps

Conversations you want to be able to check later

A doctor's appointment, a meeting with a landlord, a conversation with a bank, a family discussion about care arrangements. These are the conversations where missing one clause matters, where asking someone to repeat themselves for the fourth time is exhausting, and where you may want to reread the details a week later. A phone on the table, recording, gives you a transcript to read afterwards at your own pace, with the audio attached to every word so you can check anything that looks wrong.

Ask permission first. Recording a conversation is a normal accommodation to request, and most people say yes when you explain why, but the laws about recording differ by country and by state, and consent is the right starting point regardless of what the law requires where you are.

Lectures and classes

Lecture capture is the classic case: the room is large, the acoustics are poor, the lecturer walks around, and note-taking and lipreading compete for the same attention. Recording the session and reading the transcript afterwards separates the two tasks. A transcript is also searchable, so revising for an exam means searching for a term instead of scrubbing through six hours of audio, and topic segments make a long lecture navigable.

If your institution provides a captioner, a note-taker or lecture recordings, use those first; they are provided for a reason and they are usually better. A personal transcript is the fallback for everything the provision does not cover: the seminar, the lab session, the corridor conversation afterwards.

Meetings you are in the room for

Meetings are where automatic transcripts pay off the most, because so much of the content is procedural: who agreed to what, by when. Speaker separation matters here more than anywhere else. A wall of undifferentiated text is hard work; a transcript that shows who said each line, with names you have assigned once, reads like minutes. On top of that, a summary and a list of action items generated from the transcript answer the question you actually have, which is what am I supposed to do now, without reading forty minutes of text.

Captions for recorded media

If you are the person publishing a video, a podcast, a training module or a recorded talk, captions are what make it usable by a deaf or hard of hearing audience, and useful to everyone else. That is the other half of speech to text for deaf and hard of hearing users: one half is reading what was said to you, the other is making sure what you publish can be read. Export SRT or VTT from a transcript and you have a caption file you can attach to the video: SRT for players and editors, VTT for the web. The SRT subtitles guide covers the format and the timing fixes, and transcribing a video to text covers the import and export loop for MP4 and MOV files.

Two things worth doing before you publish captions:

  • Read them through. Automatic captions with uncorrected names, numbers and technical terms are worse than they look; a caption that quietly puts the wrong word in a speaker's mouth is a real problem, not a cosmetic one. Clicking a line to hear that exact moment makes the check quick.
  • Publish a text transcript alongside the video too. Some readers want the document rather than the captions, screen readers handle it better, and it is searchable. Markdown, PDF or Word export all work; the export guide lists the formats.

Why on-device processing matters here in particular

The conversations that most need captioning are frequently the ones you would least like to upload. Medical appointments, benefits assessments, legal advice, disciplinary meetings, family arrangements: these carry health information and other people's private words, and the other people in the room did not choose a transcription vendor.

On-device transcription means the audio and the text are processed on your Mac or iPhone, with no server in the loop, no account holding a copy and no retention policy to read. Practically, it also means it works in a hospital basement with no signal, on a plane, or in a building where the guest Wi-Fi is unusable, which is a real accessibility consideration and not an abstract one. The on-device versus cloud comparison sets out the mechanics of the distinction, including the difference between "encrypted in transit" and "never leaves the device".

There is a second reason. Access needs are personal information. Using a tool that has no account and sends nothing anywhere means your use of it is not data anyone else holds.

What this is not

Being explicit here is more respectful than being vague.

  • Not a hearing aid, and not a hearing device of any kind. It does not amplify, filter or process sound for listening. It has no medical function and makes no medical claim.
  • Not a medical device, and not a substitute for audiological care or for advice from a professional.
  • Not a real-time captioner of other people's phone calls. It does not caption calls, and it does not record or transcribe phone conversations.
  • Not a replacement for an interpreter or a professional captioner. Sign language interpretation and CART exist because automatic speech recognition is not accurate or immediate enough for legal proceedings, medical consent, education entitlements or emergencies. Where you are entitled to human provision, ask for it.
  • Not a legal record. An automatic transcript is a helpful reference, not certified evidence. Anything that has to stand up formally needs a certified transcript.
  • Not perfect. Accuracy falls with background noise, crosstalk, distance from the microphone, poor room acoustics and unfamiliar names. It is a tool that reduces effort, not one that removes the need to check things that matter.

Getting a usable transcript in practice

Automatic speech recognition rewards a small amount of preparation, and most of it is about the microphone rather than the software.

  • Put the microphone near the speech, not near you. A phone in the middle of a table beats a phone in your pocket by a wide margin, and a phone next to the person doing most of the talking beats both.
  • Record the whole session in one go. Stopping and restarting fragments the transcript and makes speaker names restart too.
  • Name the speakers once. Recognized voices are labelled automatically in later recordings, so a regular meeting gets easier over time. See teach ThinkScribe who is speaking.
  • Add the vocabulary you keep seeing wrong. Course codes, drug names, colleagues' surnames, project names: entering them once fixes them everywhere. See custom vocabulary.
  • Turn on Clean while reading. Hiding filler words and false starts makes a transcript substantially shorter and easier to read; the original text is untouched underneath.
  • Ask questions instead of reading everything. For a long session, a summary, a list of action items, or a direct question answered with citations you can click to hear is often faster than reading the full text. See the AI assistant guide.

Frequently asked questions

What is the difference between live captions and a transcript?

Live captions appear while someone is speaking, so you can follow a conversation, lecture or call as it happens; they are fast and imperfect and usually disappear afterwards. A transcript is produced from a recording and arrives after the fact, but it can be read at your own pace, searched, corrected, labelled with speaker names, summarized and exported as a document or a caption file. Many people use system live captions in the moment and a transcript afterwards.

Can I use a transcription app to caption a lecture?

You can record the lecture and read the transcript afterwards, which works well for revision because the text is searchable and every word links back to the audio. It is not live captioning, so it does not help you follow the lecture as it happens. If your school or university provides captioning, a note-taker or official recordings, use that provision first and treat a personal transcript as the fallback.

Is a transcription app a hearing aid?

No. It does not amplify or process sound for listening, it has no medical function, and it is not a medical device. It converts recorded speech into text you can read. Hearing devices and audiological care are separate things provided by professionals.

How do I add captions to a video for deaf and hard of hearing viewers?

Transcribe the video, read the transcript through and correct any names, numbers or technical terms, then export SRT or VTT. Attach the SRT next to the video for a player or import it into your editor as a caption track, and use VTT for video on a web page. Publishing a text transcript alongside the video as well helps people who prefer to read the document.

Try the private alternative.

Free to download, with free uses of every Pro feature. No account needed.

Download on the App StoreDownload on the App Store

Captions and accessibility →