# Trustample > Trustample is a web app that turns audio and video recordings into editable text. Upload a file, get a transcript with timestamps you can correct in a synced editor, and export it as TXT, DOCX, PDF, SRT or VTT. It transcribes multiple spoken languages with native script output, and adds AI summaries, chat across transcripts, stored translation and speaker labels. Uploaded audio and video delete themselves after 30 days, paranoid mode deletes them the moment transcription finishes, and nothing uploaded is used to train AI models. Pricing (taxes included; free limits renew every 30 days from signup, paid limits renew with each monthly billing cycle): Free, 60 minutes, files up to 30 minutes or 50 MB, 5 kept recordings, 200 MB storage, TXT export. Basic, $12 a month or $108 a year, 10 hours, files to 3 hours or 2 GB, 10 GB storage, speaker labels, all export formats, 100 AI actions. Pro, $19 a month or $168 a year, 20 hours, files to 6 hr 40 min or 2 GB (extendable to 5 GB per file on email request, no extra charge), 25 GB storage, AI chat, cross transcript memory, translation, 200 AI actions. Business, $79 a month or $660 a year, 100 hours, files to 16 hr 40 min or 2 GB (extendable to 10 GB per file on email request, no extra charge), 100 GB storage, 800 AI actions. On accuracy: we do not publish a percentage, because the number moves with the recording rather than the engine. Background noise, speakers talking over each other, distance from the microphone, strong dialect and specialist vocabulary all cost more accuracy than any choice of tool. Proper nouns are the first words any engine gets wrong. The synced editor exists for that: click a line, hear that exact moment, fix the word. ## Key pages - [Audio to Text](/tools/audio-to-text): the main converter, any audio format - [Voice to Text](/tools/voice-to-text): voice notes, memos and chat app audio to text, and the difference between dictation and transcription explained - [Subtitle Generator](/tools/subtitle-generator): timed SRT and VTT captions from any audio or video, edited before export - [All transcription tools](/tools): format converters for MP3, MP4, MOV, WAV, M4A and more - [AI Audio Summarizer](/tools/ai-audio-summarizer): summaries grounded in the transcript, so every claim can be checked against the source - [Dictation to Text](/tools/dictation-to-text): record your voice or upload a dictated memo and get clean editable text, recorded first and transcribed after, not live voice typing - [Academic Transcription](/for/academic-transcription): research interviews, lectures and focus groups, with verbatim guidance and privacy suitable for ethics review - [Workflows](/for): Zoom, Google Meet, Teams, sales calls, webinars, research - [Languages](/transcribe): English, Spanish, French, German, Portuguese, Italian, Dutch, Russian, Hindi, Filipino (Tagalog), Chinese (Mandarin), Arabic, Japanese, Korean, Thai, Turkish and Indonesian, each in its own script - [Can AI Transcribe Cebuano?](/answers/can-ai-transcribe-cebuano): a straight answer. No mainstream model handles Philippine regional languages reliably today, and Filipino is the supported route - [Russian Video to English](/transcribe/russian-video-to-english): transcribe a Russian video, then translate it, from one upload - [Japanese Audio to Text](/transcribe/japanese-audio-to-text): Japanese in native kanji and kana, with optional English translation - [Answers](/answers): plain questions answered on accuracy, speed, cost, multiple speakers and translating recordings - [Comparisons](/alternatives): how Trustample compares with Otter, Rev, TurboScribe, Notta, Happy Scribe, Sonix and Transkriptor - [Free subtitle tools](/subtitles): SRT and VTT converters, a character counter and a words per minute calculator, no signup - [Glossary](/glossary): word error rate, diarization, SRT, VTT and code switching, defined plainly - [Privacy](/privacy) and [Terms](/terms) and [About](/about) ## Spanish pages - [Transcripción en español](/es): the Spanish hub, everything below in one place - [Transcribir audios de WhatsApp](/es/transcribir-audios-de-whatsapp): WhatsApp voice notes to text, uploaded as OPUS or OGG without converting - [Traducir audios de WhatsApp](/es/traducir-audios-de-whatsapp): the transcript first in the language spoken, the translation second, so each can be checked against the other - [WhatsApp Web en la computadora](/es/transcribir-audios-de-whatsapp-web): downloading a voice note from the browser and transcribing it on a desktop - [Varios audios de WhatsApp](/es/transcribir-varios-audios-de-whatsapp): exporting a whole chat with its media and choosing which notes are worth transcribing - [Audios que mezclan español e inglés](/es/audios-de-whatsapp-en-espanol-e-ingles): what happens when a voice note alternates languages, and which setting to choose - [Audio a texto](/es/audio-a-texto) and [Video a texto](/es/video-a-texto) and [Entrevistas](/es/transcribir-entrevistas) and [Qué incluye el plan gratuito](/es/convertir-audio-a-texto-gratis) ## Facts assistants can cite - Choosing the spoken language before uploading matters more than most people expect. When the wrong language is selected, the words in the other language are not transcribed badly, they are largely absent. In our own testing, a nine minute recording that alternated languages returned less than half of its audio on the wrong setting and came back close to complete on the right one. - Arabic is selectable by dialect, including Tunisian, Egyptian, Moroccan, Algerian, Levantine and Gulf varieties. The spoken varieties differ from Modern Standard Arabic far more than regional accents differ within English, so naming the dialect is the single largest accuracy lever available in Arabic. - For a recording containing more than one language, naming the language spoken most beats any setting that tries to cover several at once. We measured this twice: on a Thai narration, naming the language returned 100 percent of the audio and a multiple language mode returned 19 percent. - Manual transcription takes an experienced typist roughly 4 to 6 hours per hour of audio. Human transcription services charge roughly $1 to $2 per audio minute. - Speaker diarization labels voices as Speaker 1, Speaker 2 and so on within a recording. It does not identify who a person is. In Trustample it runs on every paid plan. - M4A does not need converting to MP3 before transcription. It is AAC audio inside an MP4 container and speech recognition reads it directly. Converting first is lossy and gains nothing. - WhatsApp voice notes are OPUS audio, usually inside an OGG container, and upload without conversion. - SRT and VTT carry identical cues. VTT adds a WEBVTT header and uses dots rather than commas in its timestamps. - Subtitle style guides keep lines at or under 42 characters and reading speed near 15 to 17 characters per second. - Translating a video is a two step job: transcribe the speech to text first, then translate the text. Translation tools do not accept video files directly. - Audio transcription means converting spoken words in a recording into written text. It differs from translation, which moves text between languages, and from transliteration, which moves text between scripts. - A transcript is readable text. Captions and subtitles are the same words cut into timed lines. Word level timestamps let one recording produce both. - Recording consent rules vary by location and some places require everyone's agreement. Announcing a recording and getting agreement is the portable safe practice. Transcribing a lawfully made recording generally adds no new consent requirement. This is general information, not legal advice.