Launch offer · 31% off — KUNO €109 instead of €159 · No subscription · Designed in Munich

Kuno
EN
Buy KUNO
Product

Transcribe Audio to Text: AI Transcription Guide

Transcribe audio to text in minutes: compare free built-in tools, AI apps, on-device and human options — plus real accuracy, privacy and where your audio goes.

Published: · Reading time: ~12 min
On this page +
  1. How do I transcribe audio to text?
  2. What’s the best way to transcribe an audio file for free?
  3. How do I transcribe an audio file in Microsoft Word?
  4. How do I transcribe audio to text on iPhone and Android?
  5. Can AI transcribe audio accurately — and what does “99%” really mean?
  6. Which audio formats and languages can be transcribed?
  7. Where does your audio go when you use a transcription app?
  8. Privacy checklist before you upload sensitive audio
  9. Is it legal to transcribe a recording of someone else?
  10. Common mistakes when transcribing audio
  11. Troubleshooting: what to do when transcription fails
  12. A simple workflow that gets a clean transcript
  13. What is Kuno?

To transcribe audio to text, upload or record your file in a transcription tool and let speech-to-text AI convert it — most produce a draft in minutes. You have four realistic routes: free built-in tools you already own (Microsoft Word, Apple Voice Memos), online AI apps, on-device/offline software, or human services for near-perfect accuracy. The right one depends on your accuracy, budget and privacy needs.

💡 Quick answer

  • Fastest & free: built-in tools — Microsoft Word (“Transcribe”; availability varies by tenant, 300 min/month for uploaded audio with Microsoft 365) or Apple Voice Memos (iOS 18+, iPhone 12+; verify exact privacy/processing details before sensitive use).
  • Most accurate at scale: AI web apps (HappyScribe, Notta, Otter) ≈ 90–96% on clear audio; human review ≈ 99%.
  • Most private: on-device / offline tools — the audio never leaves your device.
  • Watch out: most “free” online tools upload your recording to a cloud, often outside the EU. Speech-to-text has moved from a niche tool to mainstream infrastructure: the global AI speech-to-text tool market is projected to grow from about $3.87 billion in 2026 to $16.42 billion by 2035 (≈17.4% CAGR, Precedence Research, verified June 2026). But faster, cheaper transcription has a quieter trade-off most guides skip — where your audio is processed and stored. This guide covers every method, what accuracy figures really mean, and how to keep sensitive recordings private.

How do I transcribe audio to text?

Every method follows the same three steps: get the audio in, run speech recognition, then clean up the text. The differences are accuracy, cost, and where the processing happens.

MethodTypical accuracy*CostWhere audio is processedBest for
Built-in OS / Office tools (Word, Apple Voice Memos, Google)85–92%Free (with subscription/device)Cloud or on-device (varies)Quick personal notes, files you already own
Online AI apps (HappyScribe, Notta, Otter, Sonix)90–96%Free tier → paidProvider cloud (often US)Interviews, podcasts, volume
On-device / offline software (Whisper, MacWhisper)90–95%Free–one-offYour own machineConfidential or offline work
Human transcription services≈99%Paid (per minute)Provider + human staffLegal, medical, publishable copy
Dedicated AI voice recorder (device)90–95%Hardware + planDevice or cloud (varies)In-person meetings, field work

*Accuracy ranges are for clear, single-speaker audio in a common language; background noise, accents and crosstalk lower every number (verified June 2026).

What’s the best way to transcribe an audio file for free?

The cheapest option is almost always a tool you already pay for or own. Three “free” routes cover most needs before you ever sign up for a dedicated app.

Free toolPlatformKey limitSupported audioProcessing
Microsoft Word “Transcribe”Word for Microsoft 365 on Windows in commercial tenants; Word for the web for government tenants300 minutes uploaded/month with Microsoft 365; 30,000 with Copilot.wav, .mp4, .m4a, .mp3Cloud / OneDrive storage
Apple Voice MemosiPhone 12+ (iOS 18+)Your own recordings onlyVoice Memos recordingsBuilt-in Apple transcript; confirm exact processing/privacy model for sensitive use
Google Docs Voice Typing / RecorderChrome, PixelLive speech, not file uploadLive microphoneCloud / on-device (Pixel)
OpenAI Whisper (open-source)Mac/Windows/LinuxTechnical setupMost common formatsYour own machine (offline)

Free online apps also exist, but the free tier is usually capped — HappyScribe currently advertises a no-credit-card free start, and its FAQ says the first 10 minutes of AI transcription are free for new users (verified June 2026). For anything recurring, treat free allowances as signup/trial offers that can change.

How do I transcribe an audio file in Microsoft Word?

Word’s built-in Transcribe tool is a fast route for a file you already have, if your Microsoft 365 tenant supports it:

  1. Open Word for Microsoft 365 on Windows in a commercial tenant, or Word for the web if you are in a government tenant, and sign in with your Microsoft 365 account.
  2. On the Home tab, open Dictate, then choose Transcribe.
  3. Select Upload audio and pick your file (.wav, .mp4, .m4a or .mp3).
  4. Wait for processing — it runs in the cloud, so time depends on file length and your connection.
  5. Review the timestamped, speaker-separated transcript, play back any section to fix errors, then Add to document.

⚠️ The 300-minute catch Microsoft Support currently states that Microsoft 365 subscribers can transcribe 300 minutes of uploaded audio per month; a Microsoft 365 Copilot license raises that to 30,000 minutes/month (verified June 2026). The support page also says feature availability differs by tenant, so check whether your account uses Word for Microsoft 365 on Windows or Word for the web. If you transcribe heavily, you’ll hit the standard limit quickly.

How do I transcribe audio to text on iPhone and Android?

On iPhone, the Voice Memos app (iOS 18 and later, iPhone 12 or newer) can generate a transcript: open a recording, tap the three-dot menu, and choose View Transcript. Apple Support confirms availability on iPhone 12 or later in English variants, Spanish, Portuguese, Italian, French, German, Japanese, Korean, Simplified Chinese and Traditional Chinese, but that support page does not state the processing model. Treat Voice Memos as a convenient built-in option; for highly sensitive recordings, verify the exact device/cloud processing path before use. On Android, Google Recorder (built into Pixel phones) transcribes live and on-device; on other Android phones, Google Docs Voice Typing or a third-party app such as Otter handles live speech. For an existing audio file on Android, you’ll generally need an app that accepts uploads, since most built-in recorders only transcribe live microphone input.

Can AI transcribe audio accurately — and what does “99%” really mean?

Yes, but read the fine print. Vendors advertise “up to 99% accuracy,” and that figure is real only for clean, single-speaker audio in a major language. HappyScribe, one of the largest providers, states its AI reaches up to 96% on clear audio, with ~99% reserved for human-reviewed transcripts (HappyScribe, verified June 2026). Accuracy is measured by Word Error Rate (WER) — the share of words inserted, deleted or substituted. Real-world factors that push WER up:

  • Background noise (cafés, traffic, air conditioning).
  • Multiple speakers and crosstalk — overlapping voices confuse speaker labels.
  • Strong accents or dialects and specialist vocabulary (names, medical or legal terms).
  • Low-quality microphones or phone-speaker recordings. Plan for editing: even a 95% transcript means roughly 1 wrong word in 20 — about a sentence per paragraph to fix.

Which audio formats and languages can be transcribed?

Most modern tools accept the common formats — MP3, WAV, M4A, AAC, FLAC, OGG — and many also extract audio from video (MP4, MOV). Leading apps support 100+ languages: HappyScribe now lists 150+ (verified June 2026), and most major engines cover English, Spanish, French, German, Italian, Portuguese and Dutch with high accuracy. Coverage thins out for less common languages and regional dialects, where accuracy drops noticeably. If your file is in an unusual format, a free audio converter or simply uploading the original video usually works — you rarely need to convert first.


Where does your audio go when you use a transcription app?

This is the question competitor pages avoid. With most online tools, your recording is uploaded to the provider’s servers — frequently in the US — to be processed, and may be retained to “improve the service.” For a podcast that’s fine. For a client call, a patient consult, an HR meeting or a confidential interview, it’s a data-protection problem.

ApproachWhere audio is processedPrivacy consideration
On-device/offline (Whisper locally, Google Recorder on Pixel, tools with explicit local processing)Your own deviceHighest — audio should not leave the device if the tool documents local/offline processing
EU-hosted AI serviceServers inside the EUStrong — data stays under EU rules
US-cloud AI serviceServers outside the EUCheck retention & whether data trains the model

Three questions to ask any tool before uploading sensitive audio: Where are the servers? How long is the audio kept? Is my data used to train AI models? If a provider can’t answer clearly, treat the recording as exposed.

Privacy checklist before you upload sensitive audio

A transcript can contain names, health details, financial information, employment context, customer secrets and exact quotes. Under GDPR, personal data is broadly any information relating to an identified or identifiable person, and even routine transcript handling counts as processing. Before uploading a recording of other people, use this checklist.

Question to checkWhy it mattersSafer default for sensitive recordings
Did everyone consent to the recording?In Germany, unauthorised recording of the non-public spoken word can trigger §201 StGB; GDPR also needs a lawful basis.Announce the purpose before recording and keep consent documented.
Where is the audio processed and stored?Cloud transcription means audio leaves your device and may leave the EU.Use on-device/offline transcription, or an EU-hosted provider with clear data-residency terms.
Is the provider a processor, and is there a DPA?For business use, the tool is usually handling personal data on your behalf.Check the data-processing agreement, sub-processors and transfer mechanism before upload.
Is audio or transcript data used for AI training?Some tools improve models from customer content unless you opt out or use an enterprise plan.Prefer explicit no-training terms, especially for HR, legal, healthcare, research or client work.
Can you delete the recording and transcript?Retention is part of the privacy risk; transcripts are more searchable than raw audio.Set a deletion deadline and confirm export/delete controls before the first sensitive upload.

This is general information, not legal advice. The practical rule is simple: if the audio would be risky to email to a stranger, do not upload it to a generic free transcription site. Use a documented business workflow or keep the transcription local/on-device.

Capture meetings without sending them to a US cloud. Kuno is a privacy-first AI voice recorder, made in Germany, built for the recordings you can’t send to an external server. It transcribes on-device, so the audio never leaves the room, is EU-hosted, and is never used to train AI. A visible recording indicator and one-tap stop also make consent clean — and it reaches the in-person and field meetings that software bots can’t. Get early access →

Transcribing audio you recorded is one thing; recording another person to transcribe them is another. In Germany, secretly recording a private spoken conversation is a criminal offence under § 201 StGB, even if you are a participant — and many countries require the consent of all parties. A transcript is also personal data under the GDPR, so you need a lawful basis (usually consent) and must store it securely. The safe rule: get clear, documented consent before recording, tell people why, and keep the file only as long as you need it. (This is general information, not legal advice; rules vary by country — see is it legal to record a conversation.)

Common mistakes when transcribing audio

  • Trusting “99% accuracy” blindly. That’s a best case on clean audio; always proofread, especially names and numbers.
  • Uploading confidential audio to a random free tool. Check where it’s processed and stored first.
  • Recording in noisy environments. Five minutes finding a quiet room beats an hour fixing errors.
  • Relying on one mic for a group. Place the device centrally or use per-speaker audio for better speaker labels.
  • Skipping speaker labels and timestamps. They make long transcripts searchable and quotable — turn them on.
  • Forgetting consent. For any recording of other people, get permission before you hit record.

Troubleshooting: what to do when transcription fails

ProblemLikely causeFix
Word “Transcribe” greyed out or missingTenant/app availability, browser/account permissions or no Microsoft 365 subscriptionUse the Microsoft-supported surface for your tenant; sign in, check Dictate > Transcribe, and confirm Edge/Chrome/mic permissions where required
”Upload limit reached”Hit the 300-minute/month capWait for reset, split files, or use another tool
Garbled or empty transcriptHeavy noise, wrong language selected, very low volumeRe-select language, denoise the file, re-record closer to the mic
Speakers not separatedDiarization off or single mixed trackEnable speaker detection; use multi-track audio if available
Unsupported file formatRare codecConvert to MP3/WAV, or upload the source video directly

A simple workflow that gets a clean transcript

  1. Record clean audio — quiet room, mic close to the speaker, one person at a time where possible.
  2. Pick the method by sensitivity — built-in/on-device for confidential audio, an AI app for volume, human review for publishable copy.
  3. Run the transcription and select the correct language up front.
  4. Proofread against the audio, fixing names, numbers and technical terms first.
  5. Export in the format you need — DOCX or PDF for documents, SRT/VTT for subtitles, TXT for raw text.

What is Kuno?

Kuno is a privacy-first AI voice recorder, made in Germany, that captures and transcribes in-person meetings on-device with EU data hosting and no training on your recordings. For transcription specifically, that means the most sensitive step — turning private speech into stored text — happens without shipping your audio to an external cloud.

Your transcript shouldn’t be the price of your privacy. Most transcription tools upload your audio to do the work. Kuno transcribes on the device, keeps data in the EU, and never trains AI on your recordings — a recorder made in Germany for meetings that have to stay confidential. Get early access →


FAQ

How can I transcribe audio to text for free? +
Use a tool you already have: Microsoft Word's Transcribe (300 uploaded minutes/month with Microsoft 365 where available), Apple Voice Memos on iPhone 12+ with iOS 18+, or Google Docs Voice Typing for live speech. For offline, open-source Whisper is free but needs setup.
What is the most accurate way to transcribe audio? +
Human transcription services reach about 99% accuracy. Among automatic tools, leading AI engines hit roughly 90–96% on clear, single-speaker audio; accuracy falls with noise, accents and multiple speakers.
Can I transcribe an audio file without uploading it to the cloud? +
Yes. Use an explicitly offline/local workflow such as Whisper running on your computer or a phone recorder that documents on-device transcription. Do not assume every built-in transcript is cloud-free; check the provider's current privacy/support page before sensitive use.
What audio formats can be transcribed? +
Most tools accept MP3, WAV, M4A, AAC, FLAC and OGG, and can extract audio from video files like MP4 and MOV. You rarely need to convert the file first.
How long does it take to transcribe audio? +
AI transcription usually takes a few minutes — text often starts appearing seconds after upload, with full files done in minutes. Human transcription typically takes hours up to 24 hours depending on length.
Is it legal to transcribe a recording of a conversation? +
Transcribing your own recordings is generally fine. Recording other people to transcribe them often requires the consent of all parties (in Germany, secret recording can breach § 201 StGB), and the transcript is personal data under the GDPR. Get consent first. *(General information, not legal advice.)*
Topics Transcription Meetings Privacy

Read next

Kuno

Stop taking notes. Connect the dots.

Kuno captures every conversation and turns it into clarity — summaries, action items, and decisions, without typing a word.

Explore Kuno