Voicely MCP Server
io.github.StulkovLD/voicely
Free, offline speech-to-text for macOS with local MCP server integration for Claude and Cursor.
What is the Voicely MCP server?
Voicely is a free, open-source offline speech-to-text application for macOS that transcribes audio and video files, records calls with speaker diarization, and enables dictation via hotkey. It exposes a local MCP server so Claude, Cursor, and other AI agents can access transcripts and perform transcription tasks entirely on-device using NVIDIA's Parakeet V3 model.
Voicely gives your Mac and AI agents the ability to transcribe audio, video, and live calls without sending data to the cloud. Press Option+Space to dictate text anywhere on your Mac, record calls with automatic speaker separation, or transcribe files through the terminal or MCP server. Everything runs locally on Apple Silicon Macs running macOS 14+, with no account, subscription, or internet required.
How to install Voicely
Copy-paste configuration for popular MCP clients.
Tools & capabilities
Tools this server exposes to the agent.
transcribe_file— Transcribe an audio or video file with optional speaker diarization.list_transcripts— List saved transcripts (dictations, calls, files) in reverse chronological order.get_transcript— Read a transcript by ID or alias (e.g., 'last', 'last-call').get_last_call— Read the most recent call transcript.
Use cases
- Dictate notes, code comments, or messages hands-free by pressing Option+Space anywhere on your Mac.
- Record Zoom, Teams, or other calls with automatic speaker identification and get a timestamped transcript.
- Transcribe interviews, podcasts, or video files from the command line or through Claude/Cursor.
- Have Claude analyze your call transcripts or dictations to extract action items, summaries, or insights.
- Integrate voice input into your AI agent workflow without relying on cloud services or API keys.
Voicely MCP server FAQ
Voicely is a free, MIT-licensed macOS app that transcribes speech to text entirely on-device using NVIDIA's Parakeet V3 model. It supports dictation via hotkey, call recording with speaker diarization, and file transcription, and exposes a local MCP server so AI agents like Claude and Cursor can access transcripts.
Yes, Voicely is completely free and open-source under the MIT license. There are no subscriptions, accounts, or cloud fees.
Install the app with `curl -fsSL https://voicely.art/install.sh | sh` or via Homebrew. Then run `voicely connect` to register the MCP server with your agents, or manually add it as a stdio server with command `voicely` and args `["mcp"]`.
No. Voicely runs entirely on-device with no cloud services, accounts, or API keys. The speech model downloads once (~470 MB) on first use and is cached locally.
Voicely supports 25 European languages including English and Russian, and can switch languages mid-sentence.
Voicely requires macOS 14 or newer on Apple Silicon (M1, M2, M3, etc.). It does not work on Intel Macs.
README (reference)
Source of truth, from the repository.
Voicely
Free, open-source, offline dictation and transcription for macOS — with a local MCP server that gives Claude Code, Codex and Cursor ears.
<p align="center"> <a href="https://github.com/StulkovLD/Voicely/actions/workflows/ci.yml"><img src="https://github.com/StulkovLD/Voicely/actions/workflows/ci.yml/badge.svg" /></a> <img src="https://img.shields.io/badge/platform-macOS%2014%2B%20%C2%B7%20Apple%20Silicon-lightgrey" /> <img src="https://img.shields.io/badge/swift-6.0-orange" /> <img src="https://img.shields.io/badge/license-MIT-green" /> <img src="https://img.shields.io/badge/engine-Parakeet%20V3%20%2B%20FluidAudio-blue" /> </p> <p align="center"><img src=".github/assets/voicely-card.png" alt="Voicely — Speak. It types. Open source, offline, free." width="600" /></p>- Dictate anywhere. Press
Option+Space, speak, press it again — the text lands wherever your cursor is, terminals and VS Code included. - Calls and files. Record a call (system audio + mic, no virtual audio device) and get a speaker-separated transcript; turn any audio or video file into text.
- Agent-native.
voicely connectregisters a local MCP server, so your agent can transcribe files and read your calls and dictations. - On-device only. NVIDIA Parakeet V3 through CoreML, 25 languages. No account, no cloud, no subscription. MIT.
curl -fsSL https://voicely.art/install.sh | sh
# or
brew install --cask stulkovld/voicely/voicely
macOS 14+ on Apple Silicon. Like Superwhisper, MacWhisper or Wispr Flow — but free and MIT-licensed, 100% on-device, and built to hand text to your AI agent.
Free, forever
Tools that do this usually ask for a monthly subscription or a one-time payment. Voicely is free, forever: MIT-licensed, and every line of it is in this repository. It is made by one person, in one room, who uses it every day. If it saves you time, a ⭐ helps others find it.
Install
One command. There is no DMG to download and nothing to drag into Applications:
curl -fsSL https://voicely.art/install.sh | sh
Or through the Homebrew tap: brew install --cask stulkovld/voicely/voicely.
The script downloads the current build, verifies its SHA-256 and its signing certificate in a private staging directory, installs it transactionally — the previous version stays as a backup until the new one passes every gate, and any failure rolls back cleanly — then launches the app. Re-running it updates in place and keeps your transcripts, model and settings.
The public build is signed with Voicely's own certificate rather than an Apple Developer ID, and it is not notarized by Apple. The certificate is the same for every release, so macOS keeps your Microphone and Accessibility permissions when you update. If an older copy left permission entries behind, the installer clears them and Voicely asks again by itself.
First launch
- Find the Voicely icon in the menu bar.
- Allow Microphone, then switch Voicely on in the Accessibility list that opens. Voicely waits for both and continues by itself.
- Confirm the speech model download (Parakeet V3, about 470 MB). The first prepare step takes about a minute; later launches are fast.
- Screen Recording is requested only when you first use Record Call.
Requirements: macOS 14 or newer on Apple Silicon.
Transport, not translation
Voicely moves what you said to where you work. It never rewrites, summarizes or "improves" what you meant — word-level truth is kept next to every transcript.
Direct transport — dictation, Option+Space
Press the hotkey, speak, press it again. A floating glass pill shows the live waveform while you talk. The text lands wherever your caret is: native Mac fields through Accessibility, checked by reading the caret back; Chrome, VS Code, Claude and every other Chromium or Electron app through the clipboard and ⌘V — your clipboard is put back right after the app has read the text, and the transcript is marked so clipboard managers don't keep it; terminals through typed text. If no app takes the text, it waits on the clipboard and the pill says "Press ⌘V". The Output menu can pin dictation to the clipboard permanently; secure fields are save-only, always. Dictations are saved to ~/Documents/Voicely/dictations/; where each one went, and how, is logged without the words in ~/Library/Logs/Voicely/insertion.jsonl.
Careful packaging — calls and files
Calls (menu bar → Record Call). Voicely records system audio (through ScreenCaptureKit — no virtual audio device needed) and your microphone, mixes them into one track, transcribes the mix and diarizes it globally. An echo can't duplicate a sentence, and several people on one mic separate into distinct voices. "You" is identified by matching diarized voices against your microphone's own activity. The transcript reads as speaker turns and is saved to ~/Documents/Voicely/calls/<id>/ as markdown, word-level JSONL and WAV.
Files. Any audio or video file becomes a transcript, with optional speaker diarization: voicely transcribe <file> from the terminal, or the transcribe_file tool from an agent. Results land in ~/Documents/Voicely/files/.
Either way the result is text in one tap. Hand it to your agent yourself — paste it into Claude Code, save it as a note — or let the agent read it through MCP, below.
Your AI knows what you know
The voicely CLI and a local MCP server make every capability agent-native. After installing the app, expose the CLI on your PATH and register the server with the agents you use:
/Applications/Voicely.app/Contents/Helpers/voicely setup # symlinks `voicely` into /usr/local/bin or ~/.local/bin, no sudo
voicely connect # registers the MCP server with every installed harness
voicely connect claude # or one by name: claude, codex, cursor, hermes, openclaw
Claude Desktop and other MCPB-aware clients can install the server in one click from voicely-<version>.mcpb on the latest release. Voicely is listed in the official MCP Registry as io.github.StulkovLD/voicely.
Teach the agent the whole playbook with the skill:
mkdir -p ~/.claude/skills/voicely && \
curl -fsSL https://voicely.art/skill/SKILL.md -o ~/.claude/skills/voicely/SKILL.md
voicely mcp speaks standard stdio MCP (JSON-RPC 2.0, fully offline) and exposes four tools:
| Tool | What it does |
|---|---|
transcribe_file | Transcribe an audio/video file (optional speaker diarization) |
list_transcripts | List saved transcripts (dictations / calls / files), newest first |
get_transcript | Read a transcript by id or alias (last, last-call, …) |
get_last_call | Read the most recent call transcript |
voicely connect uses each harness's own mcp add command and never hand-edits a config file it doesn't own. Restart the agent afterward and ask it: "transcribe ~/Downloads/interview.m4a and pull the action items", "look at my last call and draft a project plan", "what did I dictate earlier?". The skill covers bootstrap from zero, online video via yt-dlp (YouTube / TikTok / Instagram) and "watching" a video by pairing word-level timestamps with extracted frames. Everything stays on the machine.
Manual config — register voicely mcp as a stdio server. The exact file differs per harness; the shape is always command: voicely, args: ["mcp"]:
| Harness | Where | How |
|---|---|---|
| Claude Code | plugin or claude mcp add | claude mcp add voicely -- voicely mcp (or the bundled plugin) |
| Codex | ~/.codex/config.toml | [mcp_servers.voicely]<br>command = "voicely"<br>args = ["mcp"] |
| Cursor | user settings | cursor --add-mcp '{"name":"voicely","command":"voicely","args":["mcp"]}' |
| Hermes | ~/.hermes/config.yaml | hermes mcp add voicely --command voicely --args mcp |
| OpenClaw | openclaw.json | openclaw mcp set voicely --command voicely --args mcp |
| Anything else | its MCP config | {"mcpServers":{"voicely":{"command":"voicely","args":["mcp"]}}} |
There is no daemon: the server loads the speech model itself on the first transcribe_file call.
Build from source
macOS 14+ on Apple Silicon and the Xcode Command Line Tools (xcode-select --install):
git clone https://github.com/StulkovLD/Voicely.git
cd Voicely
swift build -c release --product Voicely
swift build -c release --product VoicelyCLI
.build/release/Voicely # the menu-bar app
.build/release/VoicelyCLI setup # exposes the `voicely` command
In SPM the CLI product is VoicelyCLI, not
voicely: on case-insensitive APFS avoicelyproduct would collide with the app binaryVoicely. Thevoicelycommand is created bysetup, in a directory with no such collision.
Speech model
Parakeet TDT 0.6B v3 (NVIDIA, CC-BY-4.0), running on-device through FluidAudio's CoreML port: about 470 MB on disk, 25 European languages including English and Russian — switching languages mid-sentence works — natively punctuated, roughly 30× realtime on Apple Silicon. Speaker diarization uses FluidAudio's pyannote community-1 segmentation and WeSpeaker embeddings, also on-device. Models download on first use; no Hugging Face account or token is required.
Transcripts
Everything is plain files under ~/Documents/Voicely/ — readable by you, your agent and any other tool.
Dictation (dictations/):
---
type: dictation
date: 2026-03-19T22:30:00+00:00
source_app: Telegram
---
Your transcribed text here.
Call (calls/<id>/transcript.md): one block per speaker turn; your own microphone is always You. Word-level timing lives in transcript.jsonl next to it; the audio is kept as mix.wav (what was transcribed) plus the raw mic.wav and system.wav.
---
type: call
date: 2026-08-19T18:02:11+03:00
source_app: zoom.us
duration: 12:40
speakers: You + 2
---
[00:08] You
Hey, how's the project going?
[00:11] Speaker 1
Almost done, pushing to prod tonight.
[00:17] Speaker 2
I'll review the PR after lunch.
Architecture
Option+Space ──> MicRecorder ──> Parakeet V3 (CoreML, on-device) ──> CursorInjector
│ AX insert (native) → ⌘V, clipboard restored → typed text
v │
AudioOverlay v
(liquid glass pill) TranscriptStore
(~/Documents/Voicely/…)
Record Call ──> mic + ScreenCaptureKit ──> mixdown ──> Parakeet V3 ──> FluidAudio diarization ──> speaker-turn markdown
(one track) ("You" by mic activity) + word-level JSONL + WAV
Stack: Swift 6 + AppKit + FluidAudio (Parakeet TDT v3 CoreML + diarization) + ScreenCaptureKit (system audio). All processing happens locally. No data leaves your machine.
Autostart on login
System Settings → General → Login Items → + → add Voicely.app.
Acknowledgements
- Parakeet TDT 0.6B v3 by NVIDIA (CC-BY-4.0) — the speech model.
- FluidAudio (Apache-2.0) — CoreML ports of Parakeet and the on-device diarization pipeline.
- Diarization models — pyannote segmentation and WeSpeaker embeddings (CC-BY-4.0).
- WhisperKit (MIT) — the on-device inference library this project grew up on.
License
MIT
Related MCP servers

Drive EditItAll's local in-browser editors: PDF, sheet, photo, vector, Word, slides.

SubcueAI
Public read-only MCP for SubcueAI: live pricing, latest desktop version, overview and FAQ.

Subotiz MCP
Connect AI assistants to Subotiz - Using Subotiz's external capabilities through natural language
ChatGenius tools: Instagram, Facebook, WhatsApp and SMS inbox, contacts, bookings, flows, posts.

AI memory buyer routing with Forge docs, pricing, paid recommendations, and MCP discovery.
Read-only U.S. lab-test catalog, collection-site search, and reference-range context.