Setting Up Voice Wake and Push-to-Talk
I’ve often found myself in the middle of a task, needing to send a quick message or command without wanting to break my focus. Switching windows and typing might only take a few seconds, but it interrupts the flow. I prefer being able to just speak and have the system handle the rest.
Voice Wake and Push-to-Talk are designed to solve this. They let you interact with your tools using your voice, either by waiting for a specific trigger word or by holding down a key while you speak. Here is how I set these up and what you should know about how they work.
What You’ll Need
Section titled “What You’ll Need”Before you start, make sure you have the following permissions and requirements met:
- Microphone Access: Required for capturing any audio.
- Speech Recognition Permissions: Needed for the system to turn your voice into text.
- Accessibility/Input Monitoring: Only required if you plan to use Push-to-Talk, so the app can detect the hotkey.
- macOS Version: The “Hold Cmd+Fn to talk” feature is only available on macOS 26 or later.
Quick Start
Section titled “Quick Start”You can get voice commands running in about five minutes. Follow these steps to get started:
- Enable Voice Wake: Find the Voice Wake toggle in your settings and turn it on. This starts the
VoiceWakeRuntimewhich listens for your trigger words. - Set Trigger Words: Add your preferred words to the
swabbleTriggerWordstable. The recognizer waits for these specific tokens to start capturing. - Test the Pause: When you speak, remember to leave a small gap of about 0.55s after the trigger word before you start your command. The system needs this pause to know you are starting a request.
- Configure Sounds: By default, the app uses the macOS “Glass” sound for triggers and sends. You can change this to any
NSSound-loadable file, like an MP3 or WAV, or select No Sound.
If you prefer manual control, you can also use Push-to-Talk. Hold the Right Option key (keyCode 61) to start capturing immediately. When you release the key, the transcript is finalized and sent.
How it Works
Section titled “How it Works”The VoiceWakeRuntime is always listening as long as permissions are granted. It uses specific silence windows to manage the session: it waits 2.0s if you are actively speaking, or 5.0s if it only heard the trigger word but no command followed. There is a hard stop at 120s to prevent the mic from staying on indefinitely.
When a transcript is ready, the VoiceWakeForwarder.prefixedTranscript(_:) function prepends a machine hint before sending the data to your active gateway. This ensures the receiver knows the text came from a voice session.
// Example of the forwarding logic used internallyVoiceWakeForwarder.prefixedTranscript(transcript)Troubleshooting
Section titled “Troubleshooting”If things aren’t working as expected, check these common scenarios:
- Overlay is stuck: If the overlay stays visible and the system stops listening, try clicking the “X” to dismiss it manually. This triggers a
VoiceWakeRuntime.refresh(...)via theVoiceSessionCoordinatorto resume listening. - Right Option key not working: Some external keyboards do not send the standard
keyCode 61for the Right Option key. If your keyboard is one of these, you might need to use the fallback shortcut if the app doesn’t react to the key press.
What’s Next
Section titled “What’s Next”For more help with your configuration, visit the AI Setup Assistant.
OpenClaw Expert
Still stuck?
If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.