‹ All Posts

How to Transcribe Voice Memos: Every Device, Any Length

September 14, 2026 • 9 minute read

F
Flint Team
How to Transcribe Voice Memos: Every Device, Any Length

To transcribe a voice memo, open the recording in Voice Memos on iPhone or Mac and tap the transcript button, or use Recorder on a Pixel and Voice Recorder on a Samsung. Built-in transcription is free and runs on the device. Long, noisy, or multilingual recordings need a cloud transcription engine instead.

How do I transcribe a voice memo on iPhone?

Since iOS 18, Voice Memos transcribes recordings itself, on the device, at no cost.

  1. Open Voice Memos.
  2. Tap the recording you want.
  3. Tap the transcript button, the speech-bubble icon next to the playback controls.
  4. Tap and hold the text to select and copy it.

Transcripts are searchable from the Voice Memos search field, so you can find a recording by something said inside it rather than by its date.

How do I transcribe a voice memo on Android?

Android has no single answer, because the recording app depends on who made your phone.

  • Pixel: the Recorder app transcribes live as you speak, and its transcripts are searchable. Recorder also works offline.
  • Samsung: Voice Recorder has a Speech-to-text mode. Record in that mode, or open an existing recording and tap the transcript option.
  • Any Android phone: record into Google Keep using the microphone button. Keep saves the audio and a text transcription together.

How do I transcribe a voice memo on a Mac?

Voice Memos on macOS mirrors the iPhone behaviour. Open the recording, then choose View Transcription. Recordings synced from an iPhone through iCloud appear here, so a memo captured on a walk can be transcribed on the machine where you actually write.

How do I transcribe a voice memo on Windows?

Windows has no built-in equivalent. The usual route is Word for the web, under Home, then Dictate, then Transcribe, which accepts an uploaded audio file and returns a speaker-separated transcript. Word's Transcribe needs a Microsoft 365 subscription, and Microsoft caps uploaded transcription at 300 minutes per month.

Where does built-in transcription break down?

Free built-in tools are good at one specific thing: one person, speaking clearly, in a quiet room, in the phone's own language. Push past that and quality falls off in predictable ways.

SituationWhat goes wrong
Several speakersWords are captured, but who said what is lost
Background noiseWhole phrases drop rather than degrade
Strong accentsSystematic errors on the same sounds
Technical vocabularyNames, product codes, and jargon are mangled every time
Two languages in one recordingThe engine locks onto one and garbles the other
Recordings over an hourSome tools refuse the file, others time out

Recording length catches people out most. A lecture or an interview is a different problem from a two-minute idea captured in a car, and the tool that handles the second will not always handle the first.

What should I do with recordings over an hour?

Long audio needs a cloud transcription engine, because phone-based engines trade accuracy for the ability to run locally.

Three practical options:

  • Split the file and run it through the built-in tool in chunks. Free, tedious, and speaker attribution stays broken.
  • Upload to a cloud transcription service. Better accuracy, usually priced per minute.
  • Use an app that does both, so short scraps stay on the device and long recordings go to the cloud only when it is worth it.

Flint takes the third approach. Processing Mode has three settings, and the setting can change between one recording and the next.

ModeWhere audio goesCost
On my phoneNowhere, audio stays on the deviceFree, unlimited
Credits PackFlint's transcription providerCredits per minute
Cloud, your own keyYour own Deepgram, OpenAI, Gemini, or AssemblyAI accountBilled by that provider

A common setup is on-device for everyday scraps, switched to Credits Pack for the interview that matters. See Where your audio goes for the full comparison.

On-device or cloud: which should I use?

Choosing between on-device and cloud transcription is a genuine tradeoff, not a marketing one, so both sides are worth stating bluntly.

On-device wins on privacy and cost. Nothing is uploaded, nothing is logged by a third party, there is no per-minute charge, and transcription works on a plane. In Flint's on-device mode no audio leaves the phone at all, and notes, transcripts, and audio are stored on the device rather than on a server.

Cloud wins on accuracy, and the gap is not small. Cloud engines are meaningfully better with accents, background noise, several speakers, and technical vocabulary. If a recording has to be right, the cloud is the honest recommendation.

The deciding question is not which approach is better in the abstract. The question is whether this particular recording is worth sending somewhere.

Why do the same words come out wrong every time?

A transcription engine mangles a name or a product code because the word is not in its expected vocabulary, and re-recording will not fix it. Both fixes work the same way, by telling the engine what to expect.

  • For vocabulary you use constantly, add it once to a persistent terms list. Flint calls this Settings, then Key Terms.
  • For one-off context, type it in before recording. Flint has a field on the record screen for exactly this, and anything typed there is used when the note is written up.

Can I transcribe a memo I already recorded?

Existing audio files do not need re-recording. In Flint, share the file into the app from your phone's share sheet, or open the Transcribe screen and pick the file. An export from a meeting tool, an interview recording, or a voice memo from another app all work the same way. The same screen takes a URL, so a YouTube video or a podcast episode becomes a note too. See Notes from links and files.

How do I turn a transcript into a usable note?

A raw transcript is not a note. A transcript is a wall of text with no paragraphs, every false start intact, and no indication of what mattered.

The step that makes the difference is turning that transcript into a structured write-up: a summary, the decisions, the things to do. Flint writes the note in a format you choose and keeps the raw transcript in a separate tab, so nothing is lost by cleaning it up.

Language settings are independent of all this. Input language, note language, and app language are three separate settings, so you can speak Finnish and read the note in English while the Finnish transcript stays available. Details are in Languages.

How accurate is voice memo transcription?

No transcription is perfect, and any tool claiming otherwise is selling something.

  • Clear speech, quiet room, one speaker: built-in tools are genuinely good, and a cloud engine buys you little.
  • Accents, noise, jargon, or multiple speakers: expect visible errors on-device, and reach for the cloud.
  • Long recordings: check the last thirty seconds first. If a recording was interrupted, that is where the loss lands. Flint marks recovered recordings with a Recovered badge for this reason.

One limitation worth knowing regardless of app: on iOS, recording can stop when you switch to another app or leave the screen locked for a long stretch. For a lecture, meeting, or interview, keep the recording app in the foreground.

Frequently asked questions

Can I transcribe a voice memo for free? Yes. Voice Memos on iPhone and Mac, Recorder on Pixel, and Google Keep on any Android phone all transcribe at no cost. Free built-in tools run on the device, so accuracy drops with noise, accents, and multiple speakers.

Does voice memo transcription work offline? Built-in transcription on iPhone, Mac, and Pixel runs on the device and works with no connection. Cloud transcription services need a connection, because the audio is uploaded to be processed.

Why does my transcript stop partway through the recording? A transcript that ends early usually means the recording itself ended early, not that transcription failed. On iOS, recording can stop when the app is backgrounded or the screen stays locked for a long stretch. Check the last thirty seconds of any long recording first.

How long can a voice memo be before transcription fails? Built-in tools handle short memos reliably and get unpredictable past roughly an hour, with some refusing the file and others timing out. Cloud engines handle long recordings routinely, which is the main reason to use one.

Can I transcribe a voice memo someone sent me? Yes. Save the audio file to your phone, then open it in an app that accepts file uploads. Flint takes shared audio straight from the share sheet, and Word for the web accepts uploaded files on desktop.

Does transcription work if I speak two languages in one recording? Poorly, in most tools. Engines lock onto one language and garble the other. Setting the input language explicitly, rather than relying on auto-detection, produces better results than leaving it to guess.

Is my audio uploaded anywhere when I transcribe it? On-device transcription uploads nothing. Cloud transcription uploads the audio to whichever provider does the processing. Check which mode you are in before recording anything sensitive, because the choice is made at recording time, not afterwards.

Are voice memo transcripts searchable? Yes on iPhone and Pixel, where transcripts are indexed by the system search field, so a recording can be found by a phrase spoken inside it. Transcripts produced by uploading a file to a web tool are only searchable wherever you save the resulting text.

The short answer

For a short memo on an iPhone or a Pixel, use what is already on the phone. Built-in transcription is free and runs on the device.

For a recording that is long, noisy, multilingual, or important, use a cloud engine and accept that the audio leaves the device.

To avoid choosing in advance, use something that switches per recording. For more on picking a tool, see Unlocking the Potential of AI Note Takers, or start with Getting started with Flint.