Configure OpenClaw TTS: ElevenLabs, OpenAI, or Microsoft
OpenClaw can turn your outbound replies into audio using ElevenLabs, Microsoft, or OpenAI. You can use this anywhere OpenClaw is able to send audio.
Supported services
Section titled “Supported services”- ElevenLabs (primary or fallback provider)
- Microsoft (primary or fallback provider; current bundled implementation uses
node-edge-tts) - OpenAI (primary or fallback provider; also used for summaries)
Microsoft speech notes
Section titled “Microsoft speech notes”The bundled Microsoft speech provider uses Microsoft Edge’s online neural TTS service through the node-edge-tts library. It is a hosted service that uses Microsoft endpoints and does not require an API key. While node-edge-tts provides speech configuration options and output formats, the service does not support every option. Your legacy config and directive input using edge still works and is normalized to microsoft.
Because this path is a public web service without a published SLA or quota, you should treat it as best-effort. If you need guaranteed limits and support, you should use OpenAI or ElevenLabs.
Optional keys
Section titled “Optional keys”If you want to use OpenAI or ElevenLabs, you will need these:
ELEVENLABS_API_KEY(orXI_API_KEY)OPENAI_API_KEY
Microsoft speech does not require an API key.
If you configure multiple providers, the selected provider is used first and the others act as fallback options. Auto-summary uses the configured summaryModel (or agents.defaults.model.primary), so you must ensure that provider is authenticated if you enable summaries.
Service links
Section titled “Service links”- OpenAI Text-to-Speech guide
- OpenAI Audio API reference
- ElevenLabs Text to Speech
- ElevenLabs Authentication
- node-edge-tts
- Microsoft Speech output formats
Is it enabled by default?
Section titled “Is it enabled by default?”No. Auto‑TTS is off by default. You can enable it in your config with messages.tts.auto or per session with the /tts always command (alias: /tts on).
When messages.tts.provider is unset, OpenClaw picks the first configured speech provider based on the registry auto-select order.
Config
Section titled “Config”You can find your TTS settings in openclaw.json under messages.tts. If you want to see the full technical details, take a look at the Gateway configuration.
Minimal config (enable + provider)
Section titled “Minimal config (enable + provider)”To get things running with the bare essentials, just enable the feature and pick a provider.
{ messages: { tts: { auto: "always", provider: "elevenlabs", }, },}OpenAI primary with ElevenLabs fallback
Section titled “OpenAI primary with ElevenLabs fallback”You can set up OpenAI as your main voice and keep ElevenLabs as a backup.
{ messages: { tts: { auto: "always", provider: "openai", summaryModel: "openai/gpt-4.1-mini", modelOverrides: { enabled: true, }, providers: { openai: { apiKey: "openai_api_key", baseUrl: "https://api.openai.com/v1", model: "gpt-4o-mini-tts", voice: "alloy", }, elevenlabs: { apiKey: "elevenlabs_api_key", baseUrl: "https://api.elevenlabs.io", voiceId: "voice_id", modelId: "eleven_multilingual_v2", seed: 42, applyTextNormalization: "auto", languageCode: "en", voiceSettings: { stability: 0.5, similarityBoost: 0.75, style: 0.0, useSpeakerBoost: true, speed: 1.0, }, }, }, }, },}Microsoft primary (no API key)
Section titled “Microsoft primary (no API key)”If you want to use Microsoft’s neural voices without needing an API key, this is the way to go.
{ messages: { tts: { auto: "always", provider: "microsoft", providers: { microsoft: { enabled: true, voice: "en-US-MichelleNeural", lang: "en-US", outputFormat: "audio-24khz-48kbitrate-mono-mp3", rate: "+10%", pitch: "-5%", }, }, }, },}Disable Microsoft speech
Section titled “Disable Microsoft speech”If you need to turn off the Microsoft provider specifically, use this:
{ messages: { tts: { providers: { microsoft: { enabled: false, }, }, }, },}Custom limits + prefs path
Section titled “Custom limits + prefs path”You can set a cap on
{ messages: { tts: { auto: "always", maxTextLength: 4000, timeoutMs: 30000, prefsPath: "~/.openclaw/settings/tts.json", }, },}{ messages: { tts: { auto: "inbound", }, },}{ messages: { tts: { auto: "always", }, },}/tts summary offPer-user preferences
Section titled “Per-user preferences”When you use slash commands, they write local overrides to your prefsPath. By default, this lives at ~/.openclaw/settings/tts.json, but you can override it using the OPENCLAW_TTS_PREFS environment variable or the messages.tts.prefsPath configuration.
The system stores these specific fields:
enabledprovidermaxLength(this is your summary threshold, which defaults to 1500 characters)summarize(defaults totrue)
These settings are powerful because they override the general messages.tts.* configuration for that specific host.
Output formats (fixed)
Section titled “Output formats (fixed)”The output formats are pre-configured based on the channel you are using to ensure the best compatibility.
- Feishu / Matrix / Telegram / WhatsApp: These use Opus voice messages (
opus_48000_64from ElevenLabs oropusfrom OpenAI). A 48kHz / 64kbps setting is a solid voice message tradeoff. - Other channels: These use MP3 (
mp3_44100_128from ElevenLabs ormp3from OpenAI). The 44.1kHz / 128kbps setting is the default balance for speech clarity.
If you are using Microsoft, the system uses microsoft.outputFormat, which defaults to audio-24khz-48kbitrate-mono-mp3. While the bundled transport accepts an outputFormat, you should remember that not all formats are available from the service. These values follow Microsoft Speech output formats, including Ogg/WebM Opus.
Since Telegram sendVoice accepts OGG, MP3, or M4A, you should use OpenAI or ElevenLabs if you need guaranteed Opus voice messages. If your configured Microsoft output format fails for any reason, OpenClaw retries the request with MP3.
Keep in mind that OpenAI and ElevenLabs output formats are fixed per channel as listed above.
Auto-TTS behavior
Section titled “Auto-TTS behavior”When you enable this feature, OpenClaw handles voice generation with a few smart rules to make sure it only speaks when it makes sense.
- It skips TTS if the reply already has media or a
MEDIA:directive. - It skips very short replies that are less than 10 characters.
- It summarizes long replies when you have that option enabled, using
agents.defaults.model.primary(or yoursummaryModel). - It attaches the generated audio file directly to the reply.
If a reply goes over the maxLength and you have summary turned off (or don’t have an API key for the summary model), the system skips the audio and just sends the normal text reply.
Flow diagram
Section titled “Flow diagram”Reply -> TTS enabled? no -> send text yes -> has media / MEDIA: / short? yes -> send text no -> length > limit? no -> TTS -> attach audio yes -> summary enabled? no -> send text yes -> summarize (summaryModel or agents.defaults.model.primary) -> TTS -> attach audioSlash command usage
Section titled “Slash command usage”You can manage your settings using the /tts command. If you need to enable it, take a look at the Slash commands documentation for the specifics.
Discord users should note that /tts is a built-in command on that platform. Because of this, OpenClaw registers /voice as the native command there, though typing /tts ... as text still works.
/tts off/tts always/tts inbound/tts tagged/tts status/tts provider openai/tts limit 2000/tts summary off/tts audio Hello from OpenClawKeep these points in mind:
- You need to be an authorized sender; the usual allowlist and owner rules apply.
- Make sure
commands.textor native command registration is turned on. - The options
off,always,inbound, andtaggedare per-session toggles. You can also use/tts onas a shortcut for/tts always. - Settings for
limitandsummarystay in your local preferences rather than the main config. - Use
/tts audioif you want a one-off audio reply without turning TTS on for everything. - Running
/tts statusgives you a look at how the latest attempt went. It shows success fallbacks (like<primary> -> <used>), failure errors, and detailed diagnostics including provider outcomes and latency. - If OpenAI or ElevenLabs APIs fail, you’ll see specific error details and request IDs in the logs or error messages.
Agent tool
Section titled “Agent tool”The tts tool converts text to speech and sends an audio attachment back to you. When you use Feishu, Matrix, Telegram, or WhatsApp, the system sends the audio as a voice message instead of a file attachment.
Gateway RPC
Section titled “Gateway RPC”For Gateway interactions, you can use these methods:
tts.statustts.enabletts.disabletts.converttts.setProvidertts.providers
Here you go.
[[tts:voiceId=pMsXgVXv3BLzUgSXRplE model=eleven_v3 speed=1.1]][[tts:text]](laughs) Read the song once more.[[/tts:text]]{ messages: { tts: { modelOverrides: { enabled: false, }, }, },}{ messages: { tts: { modelOverrides: { enabled: true, allowProvider: true, allowSeed: false, }, }, },}OpenClaw Expert
Still stuck?
If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.