Skip to content

Configure OpenClaw TTS: ElevenLabs, OpenAI, or Microsoft

OpenClaw can turn your outbound replies into audio using ElevenLabs, Microsoft, or OpenAI. You can use this anywhere OpenClaw is able to send audio.

  • ElevenLabs (primary or fallback provider)
  • Microsoft (primary or fallback provider; current bundled implementation uses node-edge-tts)
  • OpenAI (primary or fallback provider; also used for summaries)

The bundled Microsoft speech provider uses Microsoft Edge’s online neural TTS service through the node-edge-tts library. It is a hosted service that uses Microsoft endpoints and does not require an API key. While node-edge-tts provides speech configuration options and output formats, the service does not support every option. Your legacy config and directive input using edge still works and is normalized to microsoft.

Because this path is a public web service without a published SLA or quota, you should treat it as best-effort. If you need guaranteed limits and support, you should use OpenAI or ElevenLabs.

If you want to use OpenAI or ElevenLabs, you will need these:

  • ELEVENLABS_API_KEY (or XI_API_KEY)
  • OPENAI_API_KEY

Microsoft speech does not require an API key.

If you configure multiple providers, the selected provider is used first and the others act as fallback options. Auto-summary uses the configured summaryModel (or agents.defaults.model.primary), so you must ensure that provider is authenticated if you enable summaries.

No. Auto‑TTS is off by default. You can enable it in your config with messages.tts.auto or per session with the /tts always command (alias: /tts on).

When messages.tts.provider is unset, OpenClaw picks the first configured speech provider based on the registry auto-select order.

You can find your TTS settings in openclaw.json under messages.tts. If you want to see the full technical details, take a look at the Gateway configuration.

To get things running with the bare essentials, just enable the feature and pick a provider.

{
messages: {
tts: {
auto: "always",
provider: "elevenlabs",
},
},
}

You can set up OpenAI as your main voice and keep ElevenLabs as a backup.

{
messages: {
tts: {
auto: "always",
provider: "openai",
summaryModel: "openai/gpt-4.1-mini",
modelOverrides: {
enabled: true,
},
providers: {
openai: {
apiKey: "openai_api_key",
baseUrl: "https://api.openai.com/v1",
model: "gpt-4o-mini-tts",
voice: "alloy",
},
elevenlabs: {
apiKey: "elevenlabs_api_key",
baseUrl: "https://api.elevenlabs.io",
voiceId: "voice_id",
modelId: "eleven_multilingual_v2",
seed: 42,
applyTextNormalization: "auto",
languageCode: "en",
voiceSettings: {
stability: 0.5,
similarityBoost: 0.75,
style: 0.0,
useSpeakerBoost: true,
speed: 1.0,
},
},
},
},
},
}

If you want to use Microsoft’s neural voices without needing an API key, this is the way to go.

{
messages: {
tts: {
auto: "always",
provider: "microsoft",
providers: {
microsoft: {
enabled: true,
voice: "en-US-MichelleNeural",
lang: "en-US",
outputFormat: "audio-24khz-48kbitrate-mono-mp3",
rate: "+10%",
pitch: "-5%",
},
},
},
},
}

If you need to turn off the Microsoft provider specifically, use this:

{
messages: {
tts: {
providers: {
microsoft: {
enabled: false,
},
},
},
},
}

You can set a cap on

{
messages: {
tts: {
auto: "always",
maxTextLength: 4000,
timeoutMs: 30000,
prefsPath: "~/.openclaw/settings/tts.json",
},
},
}
{
messages: {
tts: {
auto: "inbound",
},
},
}
{
messages: {
tts: {
auto: "always",
},
},
}
/tts summary off

When you use slash commands, they write local overrides to your prefsPath. By default, this lives at ~/.openclaw/settings/tts.json, but you can override it using the OPENCLAW_TTS_PREFS environment variable or the messages.tts.prefsPath configuration.

The system stores these specific fields:

  • enabled
  • provider
  • maxLength (this is your summary threshold, which defaults to 1500 characters)
  • summarize (defaults to true)

These settings are powerful because they override the general messages.tts.* configuration for that specific host.

The output formats are pre-configured based on the channel you are using to ensure the best compatibility.

  • Feishu / Matrix / Telegram / WhatsApp: These use Opus voice messages (opus_48000_64 from ElevenLabs or opus from OpenAI). A 48kHz / 64kbps setting is a solid voice message tradeoff.
  • Other channels: These use MP3 (mp3_44100_128 from ElevenLabs or mp3 from OpenAI). The 44.1kHz / 128kbps setting is the default balance for speech clarity.

If you are using Microsoft, the system uses microsoft.outputFormat, which defaults to audio-24khz-48kbitrate-mono-mp3. While the bundled transport accepts an outputFormat, you should remember that not all formats are available from the service. These values follow Microsoft Speech output formats, including Ogg/WebM Opus.

Since Telegram sendVoice accepts OGG, MP3, or M4A, you should use OpenAI or ElevenLabs if you need guaranteed Opus voice messages. If your configured Microsoft output format fails for any reason, OpenClaw retries the request with MP3.

Keep in mind that OpenAI and ElevenLabs output formats are fixed per channel as listed above.

When you enable this feature, OpenClaw handles voice generation with a few smart rules to make sure it only speaks when it makes sense.

  • It skips TTS if the reply already has media or a MEDIA: directive.
  • It skips very short replies that are less than 10 characters.
  • It summarizes long replies when you have that option enabled, using agents.defaults.model.primary (or your summaryModel).
  • It attaches the generated audio file directly to the reply.

If a reply goes over the maxLength and you have summary turned off (or don’t have an API key for the summary model), the system skips the audio and just sends the normal text reply.

Reply -> TTS enabled?
no -> send text
yes -> has media / MEDIA: / short?
yes -> send text
no -> length > limit?
no -> TTS -> attach audio
yes -> summary enabled?
no -> send text
yes -> summarize (summaryModel or agents.defaults.model.primary)
-> TTS -> attach audio

You can manage your settings using the /tts command. If you need to enable it, take a look at the Slash commands documentation for the specifics.

Discord users should note that /tts is a built-in command on that platform. Because of this, OpenClaw registers /voice as the native command there, though typing /tts ... as text still works.

/tts off
/tts always
/tts inbound
/tts tagged
/tts status
/tts provider openai
/tts limit 2000
/tts summary off
/tts audio Hello from OpenClaw

Keep these points in mind:

  • You need to be an authorized sender; the usual allowlist and owner rules apply.
  • Make sure commands.text or native command registration is turned on.
  • The options off, always, inbound, and tagged are per-session toggles. You can also use /tts on as a shortcut for /tts always.
  • Settings for limit and summary stay in your local preferences rather than the main config.
  • Use /tts audio if you want a one-off audio reply without turning TTS on for everything.
  • Running /tts status gives you a look at how the latest attempt went. It shows success fallbacks (like <primary> -> <used>), failure errors, and detailed diagnostics including provider outcomes and latency.
  • If OpenAI or ElevenLabs APIs fail, you’ll see specific error details and request IDs in the logs or error messages.

The tts tool converts text to speech and sends an audio attachment back to you. When you use Feishu, Matrix, Telegram, or WhatsApp, the system sends the audio as a voice message instead of a file attachment.

For Gateway interactions, you can use these methods:

  • tts.status
  • tts.enable
  • tts.disable
  • tts.convert
  • tts.setProvider
  • tts.providers
Here you go.
[[tts:voiceId=pMsXgVXv3BLzUgSXRplE model=eleven_v3 speed=1.1]]
[[tts:text]](laughs) Read the song once more.[[/tts:text]]
{
messages: {
tts: {
modelOverrides: {
enabled: false,
},
},
},
}
{
messages: {
tts: {
modelOverrides: {
enabled: true,
allowProvider: true,
allowSeed: false,
},
},
},
}
OpenClaw

OpenClaw Expert

Still stuck?

If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.