Summarizing Inbound Media with OpenClaw
I have often felt the frustration of building a bot that receives a large video or a long voice memo, only to have the reply pipeline stall because the model doesn’t know what it’s looking at yet. It makes routing difficult and command parsing even harder. OpenClaw handles this by summarizing inbound media before the main reply pipeline even starts.
This “pre-digestion” phase turns media into short text descriptions. It helps the system understand what is happening in an image or what was said in an audio clip, making your bot feel much faster and more intelligent.
What You’ll Need
Section titled “What You’ll Need”Before you start, make sure you have at least one of these ready:
- Provider API Keys: OpenAI, Anthropic, Google (Gemini), Groq, or Deepgram.
- Local CLIs:
whisper-cli(whisper-cpp),whisper(Python), orsherpa-onnx-offline. - Gemini CLI: The
geminitool forread_many_filessupport.
Quick Start
Section titled “Quick Start”OpenClaw tries to be smart by auto-detecting your available tools. If you have provider keys or local CLIs on your PATH, it might already be working. If you want to take control, you can define a shared models list in your configuration.
Here is a 5-minute setup to get image and video understanding running:
{ tools: { media: { models: [ { provider: "openai", model: "gpt-5.2", capabilities: ["image"] }, { provider: "google", model: "gemini-3-flash-preview", capabilities: ["image", "audio", "video"], }, ], video: { maxChars: 500, }, }, },}When a file comes in, OpenClaw follows these steps:
- It collects the attachments.
- It picks the first model that supports that media type and fits the file size.
- It generates a summary (like
[Image] A cat sitting on a fence). - It passes this summary to your main model so it can respond immediately.
If you want to disable a specific type, like audio, just set enabled: false:
{ tools: { media: { audio: { enabled: false, }, }, },}Troubleshooting
Section titled “Troubleshooting”Sometimes things don’t go as planned. Here are two common scenarios from the logs:
- Media skipped (maxBytes): If your file is too large for a specific model (e.g., a 60MB video when the limit is 50MB), OpenClaw skips that model and tries the next one in your list. Check your
maxBytessettings if you see this often. - Audio understanding fails: If a model fails or times out, OpenClaw won’t block the reply. It simply continues the flow with the original attachments. You can check
/statusto see if a capability was skipped or if it hit a timeout.
If you aren’t sure why a CLI isn’t working, make sure the binary is on your PATH. OpenClaw expands ~ for paths, but explicit full paths in the command field are usually safer.
What’s Next
Section titled “What’s Next”OpenClaw Expert
Still stuck?
If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.