Skip to content

Enable OpenClaw Text-to-Speech: ElevenLabs & OpenAI Setup

OpenClaw converts your outbound replies into audio using providers like ElevenLabs, Microsoft, and OpenAI. This feature works in any location where OpenClaw sends audio.

You have a few choices for turning text into speech, and you can set them up as primary providers or keep them as backups.

  • ElevenLabs: Use this as your primary or fallback provider for high-quality voice.
  • Microsoft: This uses the node-edge-tts library to connect to hosted services.
  • OpenAI: This works for both speech and summaries as a primary or fallback option.
  • Fallback Logic: If you configure multiple services, OpenClaw uses your main choice first and treats the others as fallbacks.

The bundled Microsoft provider uses the Edge online neural TTS service via the node-edge-tts library. This is a hosted service that uses Microsoft endpoints, so you do not need an API key. While node-edge-tts offers various configuration options and output formats, keep in mind that the service does not support every option. If you have legacy config or directive input using edge, it still works and gets normalized to microsoft.

Since this is a public web service without a published SLA or quota, you should treat it as a best-effort option. If you need guaranteed limits and support, you are better off using OpenAI or ElevenLabs.

If you want to use OpenAI or ElevenLabs, you need to provide your keys:

  • ELEVENLABS_API_KEY (or XI_API_KEY)
  • OPENAI_API_KEY
  • Microsoft Speech: This option does not require an API key.
  • Summary Authentication: If you enable auto-summaries, the provider for your summaryModel (or agents.defaults.model.primary) must also be authenticated.

You can find more details about these services and their APIs here:

No, it is not. Auto-TTS is off by default to give you full control. You can enable it in your configuration or during a specific session:

  • Config: Use messages.tts.auto to turn it on globally.
  • Session: Use the /tts always command.
  • Alias: You can also use the /tts on alias.
  • Auto-selection: When messages.tts.provider is unset, OpenClaw picks the first configured speech provider from the registry auto-select order.

You can find your TTS settings under messages.tts in the openclaw.json file. If you need to see the full schema, check out the Gateway configuration.

If you want to get started with the basics, just set the provider and decide when audio should trigger.

{
messages: {
tts: {
auto: "always",
provider: "elevenlabs",
},
},
}

You can set up a primary provider and have another one ready as a fallback. Here is how you configure OpenAI as your main choice with ElevenLabs as the backup.

{
messages: {
tts: {
auto: "always",
provider: "openai",
summaryModel: "openai/gpt-4.1-mini",
modelOverrides: {
enabled: true,
},
providers: {
openai: {
apiKey: "openai_api_key",
baseUrl: "https://api.openai.com/v1",
model: "gpt-4o-mini-tts",
voice: "alloy",
},
elevenlabs: {
apiKey: "elevenlabs_api_key",
baseUrl: "https://api.elevenlabs.io",
voiceId: "voice_id",
modelId: "eleven_multilingual_v2",
seed: 42,
applyTextNormalization: "auto",
languageCode: "en",
voiceSettings: {
stability: 0.5,
similarityBoost: 0.75,
style: 0.0,
useSpeakerBoost: true,
speed: 1.0,
},
},
},
},
},
}

Microsoft’s option is a great choice because it works without an API key.

{
messages: {
tts: {
auto: "always",
provider: "microsoft",
providers: {
microsoft: {
enabled: true,
voice: "en-US-MichelleNeural",
lang: "en-US",
outputFormat: "audio-24khz-48kbitrate-mono-mp3",
rate: "+10%",
pitch: "-5%",
},
},
},
},
}

If you need to specifically turn off the Microsoft provider, use this configuration:

{
messages: {
tts: {
providers: {
microsoft: {
enabled: false,
},
},
},
},
}

You can control the maximum text length and change where your local preferences are stored.

{
messages: {
tts: {
auto: "always",
maxTextLength: 4000,
timeoutMs: 30000,
prefsPath: "~/.openclaw/settings/tts.json",
},
},
}

Only reply with audio after an inbound voice message

Section titled “Only reply with audio after an inbound voice message”

If you only want the system to reply with audio after you have sent a voice message, use the inbound mode.

{
messages: {
tts: {
auto: "inbound",
},
},
}

If you want to stop the system from automatically summarizing long replies, set your config to always and then use the CLI command.

{
messages: {
tts: {
auto: "always",
},
},
}

Then run:

/tts summary off
  • auto: This sets the auto-TTS mode. You can choose off, always, inbound (only after a voice message), or tagged (only when the reply has [[tts]] tags).
  • enabled: This is an old toggle that the system automatically moves to auto.
  • mode: Use "final" (default) or "all" if you want audio for tool and block replies.
  • provider: The ID for your speech provider, such as "elevenlabs", "microsoft", or "openai".
  • If you leave provider empty, OpenClaw picks the first configured provider from the registry.
  • The old provider: "edge" still works and is treated as microsoft.
  • summaryModel: An optional cheap model for summaries. It defaults to agents.defaults.model.primary.
  • modelOverrides: This allows the model to send TTS instructions. It is on by default.
  • providers.<id>: These are settings specific to each provider ID.
  • Old provider blocks like messages.tts.openai are automatically moved to messages.tts.providers.<id> when the system loads.
  • maxTextLength: The character limit for TTS input. If you go over this, /tts audio will fail.
  • timeoutMs: The request timeout in milliseconds.
  • prefsPath: Use this to override the local path for your preferences.
  • apiKey: These fall back to environment variables like ELEVENLABS_API_KEY or OPENAI_API_KEY.
  • providers.elevenlabs.baseUrl: Use this to change the ElevenLabs API endpoint.
  • providers.openai.baseUrl: Use this to change the OpenAI TTS endpoint. The system checks this config, then OPENAI_TTS_BASE_URL, then the default OpenAI URL.
  • providers.elevenlabs.voiceSettings: You can adjust stability, similarityBoost, style, useSpeakerBoost, and speed.
  • providers.elevenlabs.applyTextNormalization: Set to auto, on, or off.
  • providers.elevenlabs.languageCode: Use 2-letter ISO codes like en or de.
  • providers.elevenlabs.seed: An integer for better determinism.
  • providers.microsoft.enabled: This is true by default and requires no API key.
  • providers.microsoft.voice: The name of the Microsoft neural voice.
  • providers.microsoft.lang: The language code, such as en-US.
  • providers.microsoft.outputFormat: The audio format. Note that the Edge-backed transport might not support every format.
  • providers.microsoft.rate / pitch / volume: Use percent strings like +10%.
  • providers.microsoft.saveSubtitles: This writes JSON subtitles next to your audio file.
  • providers.microsoft.proxy: A proxy URL for these requests.
  • providers.microsoft.timeoutMs: A specific timeout for Microsoft requests.
  • edge.*: These are legacy aliases for the Microsoft settings.

By default, the model can send TTS instructions for a single reply. If you have messages.tts.auto set to tagged, these instructions are required to trigger any audio.

The model can use [[tts:...]] tags to change the voice for one reply. It can also use a [[tts:text]]...[[/tts:text]] block to include things like laughter or singing cues that should only be heard in the audio, not shown in the text.

Note that provider=... instructions are ignored unless you set modelOverrides.allowProvider: true.

Example reply payload:

Here you go.
[[tts:voiceId=pMsXgVXv3BLzUgSXRplE model=eleven_v3 speed=1.1]]
[[tts:text]](laughs) Read the song once more.[[/tts:text]]

Available directive keys:

  • provider (requires allowProvider: true)
  • voice (OpenAI) or voiceId (ElevenLabs)
  • model (OpenAI model or ElevenLabs model ID)
  • stability, similarityBoost, style, speed, useSpeakerBoost
  • applyTextNormalization
  • languageCode
  • seed

If you want to turn off all model overrides, use this:

{
messages: {
tts: {
modelOverrides: {
enabled: false,
},
},
},
}

You can also use an allowlist to let the model switch providers while keeping other settings locked:

{
messages: {
tts: {
modelOverrides: {
enabled: true,
allowProvider: true,
allowSeed: false,
},
},
},
}

You can set up personal overrides so the TTS works exactly how you want. When you use slash commands, OpenClaw saves these settings to your prefsPath. The default location is ~/.openclaw/settings/tts.json. You can change this path using the OPENCLAW_TTS_PREFS environment variable or the messages.tts.prefsPath configuration.

The system stores these specific fields:

  • enabled
  • provider
  • maxLength (this is your summary threshold, which defaults to 1500 characters)
  • summarize (defaults to true)

These local settings will override any global messages.tts.* configurations for your host.

OpenClaw uses specific audio formats depending on which chat app you use.

If you are on Feishu, Matrix, Telegram, or WhatsApp, the system sends Opus voice messages. It uses opus_48000_64 from ElevenLabs and opus from OpenAI. This 48kHz / 64kbps configuration is a great choice for voice message quality.

For other channels, the system defaults to MP3. It uses mp3_44100_128 from ElevenLabs or mp3 from OpenAI. At 44.1kHz / 128kbps, you get a clear result for speech.

Microsoft integration works a bit differently. It uses the microsoft.outputFormat setting, which defaults to audio-24khz-48kbitrate-mono-mp3. The transport accepts an outputFormat, but keep in mind that the service itself might not support every option. These values follow Microsoft Speech output formats, including Ogg and WebM Opus.

A quick tip for Telegram: while sendVoice accepts OGG, MP3, M4A, and other formats, you should use OpenAI or ElevenLabs if you need guaranteed Opus voice messages. If a configured Microsoft format fails, OpenClaw tries again with MP3.

OpenAI and ElevenLabs output formats are fixed for each channel as described above.

When you enable Auto-TTS, OpenClaw follows a specific logic to decide when to generate voice. It doesn’t just process every message; it filters them to keep the conversation natural.

OpenClaw skips TTS if the reply already contains media or a MEDIA: directive. It also skips very short replies that are less than 10 characters. For long replies, it summarizes the content using agents.defaults.model.primary (or summaryModel) and then attaches the generated audio to the reply.

If the reply exceeds your maxLength and summary is off (or you don’t have an API key for the summary model), the audio is skipped and the normal text reply is sent.

You can follow the logic path in this diagram to see exactly how OpenClaw handles your replies:

Reply -> TTS enabled?
no -> send text
yes -> has media / MEDIA: / short?
yes -> send text
no -> length > limit?
no -> TTS -> attach audio
yes -> summary enabled?
no -> send text
yes -> summarize (summaryModel or agents.defaults.model.primary)
-> TTS -> attach audio

You can manage everything with a single command: /tts. If you need to get it set up, check out the Slash commands page for all the enablement details.

If you are using Discord, keep in mind that /tts is already a built-in command there. Because of that, OpenClaw registers /voice as the native command for that platform. Don’t worry though, typing /tts ... still works exactly the same way.

/tts off
/tts always
/tts inbound
/tts tagged
/tts status
/tts provider openai
/tts limit 2000
/tts summary off
/tts audio Hello from OpenClaw

Here are a few things you should know:

  • These commands require you to be an authorized sender, so the usual allowlist and owner rules still apply.
  • You need to have commands.text or native command registration enabled for this to work.
  • The off|always|inbound|tagged options are per-session toggles. If you use /tts on, it just acts as an alias for /tts always.
  • Settings for limit and summary are saved in your local preferences, not in the main configuration file.
  • If you want a one-off audio message without turning TTS on permanently, use /tts audio.
  • When you run /tts status, you get visibility into the latest attempt, including:
    • Success fallback info: Fallback: <primary> -> <used> and Attempts: ...
    • Failure details: Error: ... and Attempts: ...
    • Detailed diagnostics: Attempt details: provider:outcome(reasonCode) latency
    • Provider specifics: OpenAI and ElevenLabs failures include parsed error details and request IDs.

The tts tool is what handles converting your text into speech. It typically returns an audio attachment so the reply can be delivered. If you are using Feishu, Matrix, Telegram, or WhatsApp, the system is smart enough to deliver the audio as a native voice message instead of a basic file attachment.

If you are working with the Gateway, you can use these specific methods:

  • tts.status
  • tts.enable
  • tts.disable
  • tts.convert
  • tts.setProvider
  • tts.providers
OpenClaw

OpenClaw Expert

Still stuck?

If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.