Skip to content

Talking to Your Assistant with Talk Mode

I find that typing out every single thought can get exhausting, especially during long coding sessions. Sometimes I just want to talk through a logic problem or ask a quick question without my hands leaving the keyboard or my eyes staying glued to the terminal.

Talk mode solves this by creating a continuous voice loop. It listens to what I say, sends it to the model, and speaks the response back to me. Here is how I set it up and how you can get it running too.

  • ElevenLabs API Key: You need this for the text-to-speech (TTS) engine.
  • Microphone and Speech Permissions: Your OS will ask for these when you first start.
  • macOS, iOS, or Android: The feature uses specific streaming playback for these platforms.

Getting Talk mode running takes about five minutes. I recommend starting with the configuration file.

Open your config file at ~/.openclaw/openclaw.json and add the talk block. Here is the standard structure:

{
"talk": {
"voiceId": "elevenlabs_voice_id",
"modelId": "eleven_v3",
"outputFormat": "mp3_44100_128",
"apiKey": "elevenlabs_api_key",
"interruptOnSpeech": true
}
}

If you don’t set a voiceId, the system defaults to your ELEVENLABS_VOICE_ID environment variable or the first voice available in your account.

On macOS, I just click Talk in the menu bar. An overlay appears to show what the assistant is doing:

  • Listening: The cloud pulses with your mic level.
  • Thinking: You will see a sinking animation.
  • Speaking: Radiating rings appear while the audio plays.

If I need to stop the assistant from talking, I click the cloud. To close the mode entirely, I click the X.

I like that the assistant can change its own voice. If the model includes a JSON line at the very beginning of its reply, OpenClaw updates the settings:

{ "voice": "new_voice_id", "once": true }

If once is true, it only changes for that specific reply. If it is missing, that voice becomes the new default for the session.

If things aren’t working as expected, check these common issues:

  • Audio isn’t playing: Ensure your outputFormat is supported. Android specifically supports pcm_16000, pcm_22050, pcm_24000, and pcm_44100.
  • Stability errors: If you use the eleven_v3 model, the stability parameter only accepts three specific values: 0.0, 0.5, or 1.0. Other models are more flexible and accept any value between 0 and 1.
  • Latency issues: On macOS or iOS, the system defaults to pcm_44100. You can try setting latency_tier to a value between 0 and 4 in your config to see if it improves response times.
  • Permission denied: Make sure you have granted both Microphone and Speech Recognition permissions in your system settings.

If you hit a wall during setup, check the AI Setup Assistant for help.

OpenClaw

OpenClaw Expert

Still stuck?

If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.