Talking to Your Assistant with Talk Mode
I find that typing out every single thought can get exhausting, especially during long coding sessions. Sometimes I just want to talk through a logic problem or ask a quick question without my hands leaving the keyboard or my eyes staying glued to the terminal.
Talk mode solves this by creating a continuous voice loop. It listens to what I say, sends it to the model, and speaks the response back to me. Here is how I set it up and how you can get it running too.
What You’ll Need
Section titled “What You’ll Need”- ElevenLabs API Key: You need this for the text-to-speech (TTS) engine.
- Microphone and Speech Permissions: Your OS will ask for these when you first start.
- macOS, iOS, or Android: The feature uses specific streaming playback for these platforms.
Quick Start
Section titled “Quick Start”Getting Talk mode running takes about five minutes. I recommend starting with the configuration file.
1. Configure OpenClaw
Section titled “1. Configure OpenClaw”Open your config file at ~/.openclaw/openclaw.json and add the talk block. Here is the standard structure:
{ "talk": { "voiceId": "elevenlabs_voice_id", "modelId": "eleven_v3", "outputFormat": "mp3_44100_128", "apiKey": "elevenlabs_api_key", "interruptOnSpeech": true }}If you don’t set a voiceId, the system defaults to your ELEVENLABS_VOICE_ID environment variable or the first voice available in your account.
2. Start Talking
Section titled “2. Start Talking”On macOS, I just click Talk in the menu bar. An overlay appears to show what the assistant is doing:
- Listening: The cloud pulses with your mic level.
- Thinking: You will see a sinking animation.
- Speaking: Radiating rings appear while the audio plays.
If I need to stop the assistant from talking, I click the cloud. To close the mode entirely, I click the X.
3. Change Voices on the Fly
Section titled “3. Change Voices on the Fly”I like that the assistant can change its own voice. If the model includes a JSON line at the very beginning of its reply, OpenClaw updates the settings:
{ "voice": "new_voice_id", "once": true }If once is true, it only changes for that specific reply. If it is missing, that voice becomes the new default for the session.
Troubleshooting
Section titled “Troubleshooting”If things aren’t working as expected, check these common issues:
- Audio isn’t playing: Ensure your
outputFormatis supported. Android specifically supportspcm_16000,pcm_22050,pcm_24000, andpcm_44100. - Stability errors: If you use the
eleven_v3model, thestabilityparameter only accepts three specific values:0.0,0.5, or1.0. Other models are more flexible and accept any value between0and1. - Latency issues: On macOS or iOS, the system defaults to
pcm_44100. You can try settinglatency_tierto a value between0and4in your config to see if it improves response times. - Permission denied: Make sure you have granted both Microphone and Speech Recognition permissions in your system settings.
If you hit a wall during setup, check the AI Setup Assistant for help.
What’s Next
Section titled “What’s Next”OpenClaw Expert
Still stuck?
If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.