openai-whisper-api
openclaw/openclaw
Transcribe audio to text via OpenAI's Whisper API with gpt-4o, mini, diarization, or whisper-1 models.
What is openai-whisper-api?
Transcribe audio files to text using OpenAI's speech-to-text API. Supports multiple models including gpt-4o-transcribe, gpt-4o-mini-transcribe, speaker diarization, and the original whisper-1. Use this when you need to convert audio recordings into searchable, editable text.
- Transcribe audio files (mp3, mp4, m4a, wav, webm, ogg, mpeg, mpga) to text
- Support multiple transcription models: gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize, and whisper-1
- Output transcriptions as plain text or JSON format
- Add speaker labels via diarization mode
- Specify language and provide context prompts to improve accuracy
- Use custom OpenAI-compatible proxies or local gateways via OPENAI_BASE_URL
How to install openai-whisper-api
npx skills add https://github.com/openclaw/openclaw --skill openai-whisper-api- curl and node installed (curl can be installed via brew)
- OPENAI_API_KEY environment variable set or configured in OpenClaw config
- Audio file in supported format (mp3, mp4, m4a, wav, webm, ogg, mpeg, mpga)
- File size under 25 MB for hosted OpenAI API
How to use openai-whisper-api
- 1.Set OPENAI_API_KEY environment variable or configure it in ~/.openclaw/openclaw.json
- 2.Run the transcribe script with your audio file: {baseDir}/scripts/transcribe.sh /path/to/audio.m4a
- 3.Optionally specify a model with --model (default is gpt-4o-transcribe)
- 4.Optionally specify output format with --out /path/to/output.txt or --json for JSON output
- 5.For speaker identification, use --model gpt-4o-transcribe-diarize
- 6.For language-specific transcription, add --language en (or other language code)
Use cases
- Transcribe meeting recordings or interviews for documentation and searchability
- Convert podcast episodes or video audio to text for accessibility and archival
- Identify and label speakers in multi-person conversations using diarization
- Batch process audio files with language-specific transcription
- Generate transcripts with custom context prompts for domain-specific terminology
- Developers integrating speech-to-text into applications
- Researchers and analysts processing audio data
- Content creators and podcasters needing transcripts
- Teams managing meeting recordings and documentation
openai-whisper-api FAQ
Supported formats include mp3, mp4, mpeg, mpga, m4a, wav, and webm. Maximum file size is 25 MB on the hosted API.
Use the --model gpt-4o-transcribe-diarize flag. This model automatically identifies and labels different speakers in the audio.
Yes, set the OPENAI_BASE_URL environment variable to point to an OpenAI-compatible proxy or local gateway.
gpt-4o-transcribe is the newer, more accurate model. whisper-1 is the original Whisper model. gpt-4o-mini-transcribe is a faster, lighter variant.
Yes, use the --prompt flag to provide speaker names or domain-specific context, except when using diarization mode which rejects prompts.
Full instructions (SKILL.md)
Source of truth, from openclaw/openclaw.
name: openai-whisper-api description: "OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1." homepage: https://platform.openai.com/docs/guides/speech-to-text metadata: { "openclaw": { "emoji": "🌐", "requires": { "bins": ["curl", "node"], "env": ["OPENAI_API_KEY"] }, "primaryEnv": "OPENAI_API_KEY", "install": [ { "id": "brew", "kind": "brew", "formula": "curl", "bins": ["curl"], "label": "Install curl (brew)", }, ], }, }
OpenAI transcriptions API
Transcribe audio through /v1/audio/transcriptions. Set OPENAI_BASE_URL for an OpenAI-compatible proxy or local gateway.
Quick start
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a
Defaults:
- Model:
gpt-4o-transcribe - Output:
<input>.txt
Useful flags
{baseDir}/scripts/transcribe.sh /path/to/audio.ogg --model gpt-4o-transcribe --out /tmp/transcript.txt
{baseDir}/scripts/transcribe.sh /path/to/audio.ogg --model gpt-4o-mini-transcribe
{baseDir}/scripts/transcribe.sh /path/to/audio.ogg --model gpt-4o-transcribe-diarize --json
{baseDir}/scripts/transcribe.sh /path/to/audio.ogg --model whisper-1
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --language en
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --prompt "Speaker names: Peter, Daniel"
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --json --out /tmp/transcript.json
Notes:
- Supported upload formats include
mp3,mp4,mpeg,mpga,m4a,wav,webm. - 25 MB upload limit on the hosted API.
- Use diarize for speaker labels; script sends
chunking_strategy=autoand rejects--prompt.
API key
Set OPENAI_API_KEY, or configure it in the active OpenClaw config file ($OPENCLAW_CONFIG_PATH, default ~/.openclaw/openclaw.json). Optionally set OPENAI_BASE_URL:
{
skills: {
"openai-whisper-api": {
apiKey: "OPENAI_KEY_HERE",
},
},
}
Related skills
More from openclaw/openclaw and the wider catalog.

openhue
Control Philips Hue lights and scenes via CLI commands.

oracle
Second-model code review and refactoring with file bundling, token preview, and API or browser execution.

ordercli
CLI for checking Foodora order history and active delivery status.

peekaboo
Capture and automate macOS UI with CLI commands for inspection, interaction, and verification.

sag
ElevenLabs text-to-speech with macOS-style say command UX.

session-logs
Search and analyze your conversation history using jq and ripgrep.