A studio microphone on a desk stand, with a cyan ring at the base of its grille.

Local speech to text: which tool to pick for meetings, dictation or files

Three open-source tools turn speech into text on your own hardware, without sending your voice to an online service, and each is built for a different job. Meetily records a meeting from the microphone and the system audio, transcribes it live, then writes up notes with a language model, local or remote. Murmure types one person's dictation into whichever window has focus, offline and on the CPU alone. Scriberr is a server that transcribes recordings you already have and labels each speaker, but it hasn't shipped a release since December 2025 and its maintainer is planning a rewrite. All three can run NVIDIA's Parakeet model; Meetily and Scriberr also run Whisper.

Meeting notes, dictation or recorded files

Meetily is built for meetings. No bot joins the call: the app records the microphone and the system audio, so the meeting software on the other end doesn't matter, and it keeps meetings and transcripts in a local SQLite database. The free edition doesn't tell speakers apart. Each segment only records its source, mic or system, and speaker separation belongs to the PRO plan, even though the GitHub repository description advertises diarization. You can also import and re-transcribe an audio file, a feature still marked beta.

Murmure is for dictation. A shortcut starts recording, and the text is typed into the active window: a document, a terminal, the prompt box of an AI assistant. Its FAQ describes it as a single-speaker tool and rules out meeting transcription. Files go through its local API, which you have to enable and which takes WAV on 127.0.0.1:4800. Since v1.11.3 the API splits audio into chunks the way dictation does, so the request size is the only limit left. The docs still mention a 5-minute cap per recording, which the changelog says silence-based chunking lifted in v1.10.1.

Scriberr starts from recordings that already exist. It's a web server written in Go that you use from a browser, and it takes audio or video files, YouTube links, multitrack Audacity projects or recordings made in the app. It returns text with word-level timestamps, exported as JSON, SRT or TXT, and separates speakers with PyAnnote, configurable up to 20 voices, or NVIDIA's Sortformer, tuned for four. A REST API calls a webhook when a job finishes, and a command-line client uploads every new file that lands in a watched folder. Its author wrote it to stop paying for the cloud transcription plan of his Plaud Note recorder, 100 dollars a year for 20 hours a month.

Whisper or Parakeet: which models run where

Whisper runs in Meetily and Scriberr. Meetily uses whisper.cpp, with models from tiny to large-v3 downloaded from Hugging Face and large-v3-turbo as the default. Scriberr uses WhisperX, also from tiny to large-v3, which covers more than 90 languages and can translate into English.

NVIDIA's Parakeet TDT 0.6B v3 is the only model in Murmure, which runs it on the CPU. The project's site says it matches whisper-large's accuracy with 0.6 billion parameters against 1.5 billion. Meetily ships it too, as an int8-quantized ONNX model.

The same model doesn't cover the same languages from one tool to the next. Murmure uses it for 25 European languages and has no setting to force one: with a poor microphone or background noise, the output falls back to English, which its docs call the most commonly reported issue. French, English, German and Swedish are recognized best. Scriberr's adapter code registers Parakeet v3 as English-only, so anything else goes to WhisperX or Canary 1B v2, another NVIDIA model.

What each tool needs to run

Read from the repositories and documentation on 7 October 2026.

ToolLicenceForm and platformsTranscription modelsDocumented hardwareLatest release
MeetilyMIT, Community Editiondesktop app, SQLite; installers for Windows x64 and Apple Silicon Macs, build from source elsewhereWhisper (whisper.cpp), Parakeet TDT 0.6B v3AVX2 CPU on Windows; GPU through Vulkan or Metal, CUDA if you compile; memory not documented0.4.1, 12 Sep 2026
MurmureAGPL-3.0 or laterdesktop app, no database; Windows 10+, Apple Silicon and Intel Macs, Linux as .deb, .rpm, pacman or AppImageParakeet TDT 0.6B v3CPU only, 2 GB of free RAM, 1 GB of disk1.11.3, 31 Aug 2026
ScriberrMIT for the code, CC-BY-4.0 for the NVIDIA modelsGo server, SQLite, Python environments; Docker or Homebrew; single userWhisperX, Canary 1B v2, Parakeet v3 (English); diarization with PyAnnote or SortformerCPU or NVIDIA GPU; 2 GB (WhisperX) to 8 GB and up (Canary) per the code1.2.0, 17 Dec 2025

Murmure puts numbers on its needs: its site asks for 2 GB of free RAM and 1 GB of disk, with no graphics card. The v1.11.3 installers weigh 600 to 865 MB. Prompt Mode, which hands the dictation to an LLM to translate or rewrite it, asks for more: without a GPU it is very slow, and the docs suggest 4 to 8 GB of VRAM depending on the model. On Linux, the .deb needs GLIBC 2.38, as shipped with Ubuntu 24.04, and on Wayland Murmure registers no global shortcut: you bind a desktop shortcut to murmure --transcription, with no push-to-talk.

Meetily's README gives no memory figure for transcription. Its installers cover Windows x64, with an AVX2-capable CPU and a Vulkan build of Whisper, and macOS on Apple Silicon, with Metal and Core ML. On Linux or an Intel Mac you build from source with Rust, Node.js and pnpm, and CUDA also means compiling it yourself. Local summaries run through llama.cpp, bundled in the app, with four GGUF models; the app lists Gemma 3 1B at about 1 GB of RAM and Gemma 3 4B at about 3.5 GB.

Scriberr's docs give no memory figures, but its code recommends 2 GB for WhisperX, 4 GB for Parakeet and 8 GB or more for Canary, where a GPU is "strongly" recommended. Only NVIDIA cards speed up transcription: the docs provide a CPU Docker image, a CUDA image for GTX 10 through RTX 40 cards and a separate one for the RTX 50 series, and nothing for AMD. On macOS, Parakeet, Canary, PyAnnote and Sortformer install the CPU build of PyTorch. The first start sets up one Python environment per model family and downloads several gigabytes: 2.3 GB for Parakeet and 5.9 GB for Canary in the sample log from the docs. Because its production cookies are flagged Secure, Scriberr served over plain HTTP fails to load audio unless you set SECURE_COOKIES=false.

What leaves your machine

By default, Meetily, Murmure and Scriberr all keep the audio on your machine. The differences come from the services you plug into them.

Meetily sends the transcript text, never the audio, to the summary provider when you pick a remote one: Claude, Groq, OpenRouter, OpenAI or any OpenAI-compatible endpoint. To keep summaries local, use the bundled llama.cpp or Ollama. The app downloads its models from Hugging Face, and Parakeet from meetily.towardsgeneralintelligence.com. Its PostHog analytics have been off by default since v0.4.0; when turned on, they send usage metrics and no meeting content.

Murmure checks GitHub for updates on every launch, and the app has no switch to stop it; according to its privacy policy, the request carries no audio or text. Pointing Prompt Mode at a remote server sends your transcriptions there, and Murmure warns about it during setup; the same mode also works with a local Ollama. History stays in memory and is wiped when the app closes, unless you turn on the option that keeps the last five transcriptions on disk.

Scriberr sends nothing out by default. Two options do send content elsewhere: transcription through the OpenAI API, inactive while OPENAI_API_KEY is empty, and summaries or questions handled by a remote OpenAI-compatible API instead of Ollama. PyAnnote diarization needs a Hugging Face account, accepted model terms and an access token, which is used to download the models; according to the docs, no audio is sent.

Where the projects stand

Meetily is at v0.4.1, released on 12 September 2026 and mostly bug fixes, its fourth release of the year. On 7 October 2026 its repository had 202 open issues and 158 open pull requests.

Murmure's repository dates from October 2025, and the project shipped a minor release every one to two months through 2026, up to v1.11.3 on 31 August. One developer authored 554 of the 603 commits on the main branch.

Scriberr hasn't shipped anything since v1.2.0 on 17 December 2025. Later commits, through March 2026, add Mistral's Voxtral Mini among other things, and no release includes it. The maintainer, who wrote 636 of the 828 commits, paused development in March 2026, then said he was back on 20 September in a GitHub discussion on the project's future. In it he leans toward rebuilding the engine on onnx-asr or audio.cpp rather than working through the open issues and 21 open pull requests, and expects the rewrite would "probably" mean breaking changes. Upgrading from v1.1 to v1.2 already forced one: deleting the old Python environment and splitting the volumes.

Licences and paid plans

Meetily comes as two products. The Community Edition is under MIT, and its README promises it will stay "free & open source forever". PRO, at 10 dollars per user per month billed yearly, is sold under a commercial licence and built on a separate codebase. Its sales page lists speaker separation, live or on imported audio, and the README adds custom summary templates, PDF, DOCX and Markdown exports, automatic meeting detection and self-hosted deployment for teams of 2 to 100 people. Larger organisations are pointed to an Enterprise tier with custom pricing. The README still marks speaker identification as "Coming Soon", while the PRO page presents it as available.

Murmure is under AGPL-3.0 or later, so derivative works must stay open source. It has no paid tier ("No business model," says its site) and takes donations through Tipeee.

Scriberr is under MIT, but that licence only covers its code: Parakeet, Canary and Sortformer are under CC-BY-4.0, and PyAnnote's models require accepting their terms on Hugging Face. There is no paid plan and no cloud service; the project takes donations on Ko-fi and lists Recall.ai as a sponsor.