Skip to content

How to Handle Voice Notes and Audio Transcription

I often find myself in situations where I cannot listen to a voice note. Maybe I am in a loud room or a quiet meeting, but I still need to know what was said so I can react. When you are building tools, handling these audio files manually is a pain. You want the system to just “get it” and turn that sound into text you can actually use.

OpenClaw handles this by locating audio attachments and running them through a transcription process. It replaces the message body with the transcript, which means your slash commands still work even if you sent them as a voice memo. I find this approach much better than manually downloading and uploading files to a separate service.

  • An audio attachment (either a local path or a URL).
  • tools.media.audio.enabled set to true (this is the default).
  • A configured provider key (OpenAI, Groq, Deepgram, or Google).
  • Local CLI tools like whisper or whisper-cpp if you want to process audio locally.

OpenClaw uses auto-detection by default. If you do not configure specific models, it will try to find a way to transcribe your audio in this order:

  1. Local CLIs: It looks for sherpa-onnx-offline, whisper-cli, or the standard Python whisper.
  2. Gemini CLI: It tries the gemini command using read_many_files.
  3. Cloud Providers: It checks for OpenAI, Groq, Deepgram, and then Google.

If you want to be specific, I recommend setting up a provider with a local fallback. Here is how you can set up OpenAI with a local Whisper fallback in your config:

{
tools: {
media: {
audio: {
enabled: true,
maxBytes: 20971520,
models: [
{ provider: "openai", model: "gpt-4o-mini-transcribe" },
{
type: "cli",
command: "whisper",
args: ["--model", "base", "{{MediaPath}}"],
timeoutSeconds: 45,
},
],
},
},
},
}

Once this is set, OpenClaw will download the audio, check if it is under the 20MB limit, and run the first model that works. The resulting text is then available in your templates as {{Transcript}}.

If you prefer using Deepgram, the setup is quite simple. It will automatically use your DEEPGRAM_API_KEY.

{
tools: {
media: {
audio: {
enabled: true,
models: [{ provider: "deepgram", model: "nova-3" }],
},
},
},
}
  • Audio file is skipped: Check the file size. The default limit is 20MB (maxBytes: 20971520). If a file is too large, OpenClaw skips that model and tries the next one in your list.
  • CLI output issues: Ensure your local CLI prints plain text and exits with code 0. If your CLI tool returns JSON, you will need to pipe it through jq -r .text to get the raw transcript.
  • Slow responses: Transcription can take time. Check your timeoutSeconds (the default is 60s). If the process takes too long, it might block the reply queue.
  • Too much noise in groups: If you don’t want audio processed in group chats, use scope rules. The following config denies audio processing in groups but allows it everywhere else:
{
tools: {
media: {
audio: {
enabled: true,
scope: {
default: "allow",
rules: [{ action: "deny", match: { chatType: "group" } }],
},
models: [{ provider: "openai", model: "gpt-4o-mini-transcribe" }],
},
},
},
}

If you need help with your specific configuration, check out the AI Setup Assistant.

OpenClaw

OpenClaw Expert

Still stuck?

If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.