dev.voicemode/voicemode MCP Server
dev.voicemode/voicemode
Natural voice conversations with Claude Code—speak naturally, hear responses immediately with STT/TTS via MCP.
What is the dev.voicemode/voicemode MCP server?
VoiceMode is an MCP server that enables natural voice conversations with Claude Code and other MCP-capable agents through speech-to-text and text-to-speech. It supports both cloud-based services (OpenAI) and local offline alternatives (Whisper.cpp, Kokoro) for privacy-conscious users. Perfect for hands-free interaction when typing or looking at screens isn't practical.
VoiceMode lets you have natural voice conversations with Claude Code using your microphone and speakers. It handles speech-to-text transcription and text-to-speech responses with smart silence detection and low latency. You can run it entirely offline with local voice services or use cloud APIs—it works seamlessly across Linux, macOS, Windows, and NixOS.
How to install dev.voicemode/voicemode
Copy-paste configuration for popular MCP clients.
OPENAI_API_KEYsecretOpenAI API key for cloud-based STT/TTS (optional - local services can be installed)
VOICEMODE_DEBUGEnable debug mode with detailed logging (true/false)
VOICEMODE_SKIP_TTSSkip TTS and show text only for faster response (true/false)
VOICEMODE_PREFER_LOCALPrefer local services over cloud when available (true/false, default: true)
VOICEMODE_AUDIO_FORMATAudio format: pcm, mp3, wav, flac, aac, opus (default: pcm)
VOICEMODE_WHISPER_MODELWhisper model: tiny, base, small, medium, large (default: base)
VOICEMODE_DISABLE_SILENCE_DETECTIONDisable silence detection for continuous recording (true/false)
Tools & capabilities
Tools this server exposes to the agent.
converse— Start a natural voice conversation with Claude Code, with automatic speech-to-text transcription and text-to-speech responsesservice— Manage local voice services (Whisper.cpp for STT, Kokoro for TTS) or configure cloud service integration
Use cases
- Have hands-free conversations with Claude while walking, cooking, or doing other activities
- Debug code and get assistance without typing, using only voice commands and audio responses
- Use local speech services to keep conversations completely private and offline
- Switch seamlessly between cloud-based and local voice services depending on your privacy needs
- Get immediate audio feedback from Claude Code with natural conversation flow and smart silence detection
dev.voicemode/voicemode MCP server FAQ
VoiceMode is an MCP server that adds natural voice conversation capabilities to Claude Code. You speak into your microphone, it transcribes your speech, sends it to Claude, and reads Claude's response aloud—all with low latency and smart silence detection.
VoiceMode itself is free and open-source (MIT license). If you use cloud services like OpenAI's Whisper or TTS, you'll incur API costs. Local alternatives (Whisper.cpp, Kokoro) are free and run entirely on your machine.
The fastest way is via the Claude Code plugin marketplace: `claude plugin marketplace add mbailey/voicemode`, then `claude plugin install voicemode@voicemode`. Alternatively, install the Python package (`uvx voice-mode-install`) and add it as an MCP server.
Yes. VoiceMode supports local speech services: Whisper.cpp for speech-to-text and Kokoro for text-to-speech. Both run entirely on your machine without requiring internet or API keys.
VoiceMode runs on Linux, macOS, Windows (native or WSL), and NixOS. It requires Python 3.10-3.14 and a computer with a microphone and speakers.
VoiceMode works out of the box. For seamless operation without permission prompts, you can optionally add permissions to `~/.claude/settings.json` to allow the `mcp__voicemode__converse` and `mcp__voicemode__service` tools.
README (reference)
Source of truth, from the repository.
VoiceMode
Natural voice conversations with Claude Code (and other MCP capable agents)
VoiceMode enables natural voice conversations with Claude Code. Voice isn't about replacing typing - it's about being available when typing isn't.
Perfect for:
- Walking to your next meeting
- Cooking while debugging
- Giving your eyes a break after hours of screen time
- Holding a coffee (or a dog)
- Any moment when your hands or eyes are busy
See It In Action
Quick Start
Requirements: Computer with microphone and speakers
Option 1: Claude Code Plugin (Recommended)
The fastest way for Claude Code users to get started:
# Add the VoiceMode marketplace
claude plugin marketplace add mbailey/voicemode
# Install VoiceMode plugin
claude plugin install voicemode@voicemode
## Install dependencies (CLI, Local Voice Services)
/voicemode:install
# Start talking!
/voicemode:converse
Option 2: Python installer package
Installs dependencies and the VoiceMode Python package.
# Install UV package manager (if needed)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Run the installer (sets up dependencies and local voice services)
uvx voice-mode-install
# Add to Claude Code
claude mcp add --scope user voicemode -- uvx --refresh --from voice-mode voicemode-mcp-launcher
# Optional: Add OpenAI API key as fallback for local services
export OPENAI_API_KEY=your-openai-key
# Start a conversation
claude converse
For manual setup, see the Getting Started Guide.
Features
- Natural conversations - speak naturally, hear responses immediately
- Works offline - optional local voice services (Whisper STT, Kokoro TTS)
- Low latency - fast enough to feel like a real conversation
- Smart silence detection - stops recording when you stop speaking
- Privacy options - run entirely locally or use cloud services
Compatibility
Platforms: Linux, macOS, Windows (native or WSL), NixOS Python: 3.10-3.14
Configuration
VoiceMode works out of the box. For customization:
# Set OpenAI API key (if using cloud services)
export OPENAI_API_KEY="your-key"
# Or configure via file
voicemode config edit
See the Configuration Guide for all options.
Permissions Setup (Optional)
To use VoiceMode without permission prompts, add to ~/.claude/settings.json:
{
"permissions": {
"allow": [
"mcp__voicemode__converse",
"mcp__voicemode__service"
]
}
}
See the Permissions Guide for more options.
Local Voice Services
For privacy or offline use, install local speech services:
- Whisper.cpp - Local speech-to-text
- Kokoro - Local text-to-speech with multiple voices
These provide the same API as OpenAI, so VoiceMode switches seamlessly between them.
Installation Details
<details> <summary><strong>System Dependencies by Platform</strong></summary>Ubuntu/Debian
sudo apt update
sudo apt install -y ffmpeg gcc libasound2-dev libasound2-plugins libportaudio2 portaudio19-dev pulseaudio pulseaudio-utils python3-dev
WSL2 users: The pulseaudio packages above are required for microphone access.
Fedora/RHEL
sudo dnf install alsa-lib-devel ffmpeg gcc portaudio portaudio-devel python3-devel
macOS
brew install ffmpeg node portaudio
NixOS
# Use development shell
nix develop github:mbailey/voicemode
# Or install system-wide
nix profile install github:mbailey/voicemode
</details>
<details>
<summary><strong>Alternative Installation Methods</strong></summary>
From source
git clone https://github.com/mbailey/voicemode.git
cd voicemode
uv tool install -e .
NixOS system-wide
# In /etc/nixos/configuration.nix
environment.systemPackages = [
(builtins.getFlake "github:mbailey/voicemode").packages.${pkgs.system}.default
];
</details>
Troubleshooting
| Problem | Solution |
|---|---|
| No microphone access | Check terminal/app permissions. WSL2 needs pulseaudio packages. |
| UV not found | Run curl -LsSf https://astral.sh/uv/install.sh | sh |
| OpenAI API error | Verify OPENAI_API_KEY is set correctly |
| No audio output | Check system audio settings and available devices |
Save Audio for Debugging
export VOICEMODE_SAVE_AUDIO=true
# Files saved to ~/.voicemode/audio/YYYY/MM/
Documentation
- Getting Started - Full setup guide
- Configuration - All environment variables
- Whisper Setup - Local speech-to-text
- Kokoro Setup - Local text-to-speech
- Development Setup - Contributing guide
Full documentation: voicemode.dev
Links
- Website: voicemode.dev
- GitHub: github.com/mbailey/voicemode
- PyPI: pypi.org/project/voice-mode
- YouTube: @getvoicemode
- Twitter/X: @getvoicemode
- Newsletter:
License
MIT - A Failmode Project
mcp-name: dev.voicemode/voicemode
Related MCP servers
Read-only production error evidence for coding agents investigating failures.
Create, update, and revoke Apple Wallet and Google Wallet passes from any MCP client.

dev.waxberry/live-translate-mcp
MCP server for local speech translation (EN ↔ 中文) via Whisper + Claude + Piper
Audit Solidity and Rust smart contracts from your editor. Pays per audit in USDC over x402.

Free website analyzer: score any public URL 0-100 across 8 quality dimensions. No auth.

WebLens
Scrape, crawl, map and extract the web. Pay per call in USDC, no account or API key.
