How to Handle Voice Notes and Audio Transcription
I often find myself in situations where I cannot listen to a voice note. Maybe I am in a loud room or a quiet meeting, but I still need to know what was said so I can react. When you are building tools, handling these audio files manually is a pain. You want the system to just “get it” and turn that sound into text you can actually use.
OpenClaw handles this by locating audio attachments and running them through a transcription process. It replaces the message body with the transcript, which means your slash commands still work even if you sent them as a voice memo. I find this approach much better than manually downloading and uploading files to a separate service.
What You’ll Need
Section titled “What You’ll Need”- An audio attachment (either a local path or a URL).
tools.media.audio.enabledset totrue(this is the default).- A configured provider key (OpenAI, Groq, Deepgram, or Google).
- Local CLI tools like
whisperorwhisper-cppif you want to process audio locally.
Quick Start
Section titled “Quick Start”OpenClaw uses auto-detection by default. If you do not configure specific models, it will try to find a way to transcribe your audio in this order:
- Local CLIs: It looks for
sherpa-onnx-offline,whisper-cli, or the standard Pythonwhisper. - Gemini CLI: It tries the
geminicommand usingread_many_files. - Cloud Providers: It checks for OpenAI, Groq, Deepgram, and then Google.
If you want to be specific, I recommend setting up a provider with a local fallback. Here is how you can set up OpenAI with a local Whisper fallback in your config:
{ tools: { media: { audio: { enabled: true, maxBytes: 20971520, models: [ { provider: "openai", model: "gpt-4o-mini-transcribe" }, { type: "cli", command: "whisper", args: ["--model", "base", "{{MediaPath}}"], timeoutSeconds: 45, }, ], }, }, },}Once this is set, OpenClaw will download the audio, check if it is under the 20MB limit, and run the first model that works. The resulting text is then available in your templates as {{Transcript}}.
Deepgram Setup
Section titled “Deepgram Setup”If you prefer using Deepgram, the setup is quite simple. It will automatically use your DEEPGRAM_API_KEY.
{ tools: { media: { audio: { enabled: true, models: [{ provider: "deepgram", model: "nova-3" }], }, }, },}Troubleshooting
Section titled “Troubleshooting”- Audio file is skipped: Check the file size. The default limit is 20MB (
maxBytes: 20971520). If a file is too large, OpenClaw skips that model and tries the next one in your list. - CLI output issues: Ensure your local CLI prints plain text and exits with code 0. If your CLI tool returns JSON, you will need to pipe it through
jq -r .textto get the raw transcript. - Slow responses: Transcription can take time. Check your
timeoutSeconds(the default is 60s). If the process takes too long, it might block the reply queue. - Too much noise in groups: If you don’t want audio processed in group chats, use scope rules. The following config denies audio processing in groups but allows it everywhere else:
{ tools: { media: { audio: { enabled: true, scope: { default: "allow", rules: [{ action: "deny", match: { chatType: "group" } }], }, models: [{ provider: "openai", model: "gpt-4o-mini-transcribe" }], }, }, },}If you need help with your specific configuration, check out the AI Setup Assistant.
What’s Next
Section titled “What’s Next”OpenClaw Expert
Still stuck?
If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.