Setting up Deepgram for Audio Transcription
I’ve found that handling audio files can be one of the most annoying parts of building any application. You receive a voice note and just want the text, but setting up the pipeline to process it usually feels like extra work. If you are using OpenClaw, I recommend using Deepgram to handle your inbound audio transcription.
It is a speech-to-text API that OpenClaw uses for voice notes via tools.media.audio. When you turn it on, OpenClaw uploads the file to Deepgram and puts the transcript right into the reply pipeline using the {{Transcript}} and [Audio] blocks. This uses the pre-recorded transcription endpoint, so keep in mind it is not for streaming.
What You’ll Need
Section titled “What You’ll Need”- A Deepgram API key from deepgram.com
- OpenClaw installed and configured
Quick Start
Section titled “Quick Start”I’ve found the easiest way to get this running is to follow these steps. You can have it working in about five minutes.
-
Set your API key Add your key to your environment variables. This is the simplest path for authentication.
DEEPGRAM_API_KEY=dg_... -
Enable the provider Update your configuration file to enable the audio tool and set Deepgram as the provider.
{tools: {media: {audio: {enabled: true,models: [{ provider: "deepgram", model: "nova-3" }],},},},} -
Add a language hint (Optional) If you know the language being spoken, you can specify it in the model configuration.
{tools: {media: {audio: {enabled: true,models: [{ provider: "deepgram", model: "nova-3", language: "en" }],},},},} -
Configure advanced options If you want features like punctuation or smart formatting, use the
providerOptionsblock.{tools: {media: {audio: {enabled: true,providerOptions: {deepgram: {detect_language: true,punctuate: true,smart_format: true,},},models: [{ provider: "deepgram", model: "nova-3" }],},},},}
Troubleshooting
Section titled “Troubleshooting”If things aren’t working as expected, check these common areas:
- Authentication Errors: Verify that
DEEPGRAM_API_KEYis set correctly. OpenClaw follows a standard provider auth order, and this variable is the most direct method. - Proxy Issues: If you are using a proxy, you might need to override the default settings. Use
tools.media.audio.baseUrlandtools.media.audio.headersto point to your proxy. - Missing Transcripts: Check if the file exceeds size caps or if the request is hitting timeouts.
- Output Rules: Deepgram follows the same audio rules as other providers regarding transcript injection and size limits.
If you need more help with your specific configuration, check out the AI Setup Assistant.
What’s Next
Section titled “What’s Next”OpenClaw Expert
Still stuck?
If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.