Skip to content

Setting Up Voice Wake and Push-to-Talk

I’ve often found myself in the middle of a task, needing to send a quick message or command without wanting to break my focus. Switching windows and typing might only take a few seconds, but it interrupts the flow. I prefer being able to just speak and have the system handle the rest.

Voice Wake and Push-to-Talk are designed to solve this. They let you interact with your tools using your voice, either by waiting for a specific trigger word or by holding down a key while you speak. Here is how I set these up and what you should know about how they work.

Before you start, make sure you have the following permissions and requirements met:

  • Microphone Access: Required for capturing any audio.
  • Speech Recognition Permissions: Needed for the system to turn your voice into text.
  • Accessibility/Input Monitoring: Only required if you plan to use Push-to-Talk, so the app can detect the hotkey.
  • macOS Version: The “Hold Cmd+Fn to talk” feature is only available on macOS 26 or later.

You can get voice commands running in about five minutes. Follow these steps to get started:

  1. Enable Voice Wake: Find the Voice Wake toggle in your settings and turn it on. This starts the VoiceWakeRuntime which listens for your trigger words.
  2. Set Trigger Words: Add your preferred words to the swabbleTriggerWords table. The recognizer waits for these specific tokens to start capturing.
  3. Test the Pause: When you speak, remember to leave a small gap of about 0.55s after the trigger word before you start your command. The system needs this pause to know you are starting a request.
  4. Configure Sounds: By default, the app uses the macOS “Glass” sound for triggers and sends. You can change this to any NSSound-loadable file, like an MP3 or WAV, or select No Sound.

If you prefer manual control, you can also use Push-to-Talk. Hold the Right Option key (keyCode 61) to start capturing immediately. When you release the key, the transcript is finalized and sent.

The VoiceWakeRuntime is always listening as long as permissions are granted. It uses specific silence windows to manage the session: it waits 2.0s if you are actively speaking, or 5.0s if it only heard the trigger word but no command followed. There is a hard stop at 120s to prevent the mic from staying on indefinitely.

When a transcript is ready, the VoiceWakeForwarder.prefixedTranscript(_:) function prepends a machine hint before sending the data to your active gateway. This ensures the receiver knows the text came from a voice session.

// Example of the forwarding logic used internally
VoiceWakeForwarder.prefixedTranscript(transcript)

If things aren’t working as expected, check these common scenarios:

  • Overlay is stuck: If the overlay stays visible and the system stops listening, try clicking the “X” to dismiss it manually. This triggers a VoiceWakeRuntime.refresh(...) via the VoiceSessionCoordinator to resume listening.
  • Right Option key not working: Some external keyboards do not send the standard keyCode 61 for the Right Option key. If your keyboard is one of these, you might need to use the fallback shortcut if the app doesn’t react to the key press.

For more help with your configuration, visit the AI Setup Assistant.

OpenClaw

OpenClaw Expert

Still stuck?

If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.