Skip to content

Keep Your Context Lean with Session Pruning

I’ve spent a lot of time watching chat sessions grow until they become slow and expensive. When you work with tools that return large amounts of data, your context window fills up fast with information that the model might not need anymore. If you use prompt caching, an idle session can lead to high costs because you have to re-cache the entire history once the cache expires.

I prefer a cleaner approach. Session pruning helps by trimming old tool results from the in-memory context right before each call to the model. It keeps your history on disk exactly as it is but sends a leaner version to the API.

  • Anthropic API or OpenRouter Anthropic models.

Pruning is turned off by default. You can enable it by setting the mode to cache-ttl in your configuration. This ensures pruning only happens when your session has been idle longer than the cache window.

To get started, add this to your configuration:

{
agent: {
contextPruning: { mode: "cache-ttl", ttl: "5m" },
},
}

If you want to limit pruning to specific tools, you can use an allow list:

{
agent: {
contextPruning: {
mode: "cache-ttl",
tools: { allow: ["exec", "read"], deny: ["*image*"] },
},
},
}

I find it helpful to understand exactly what stays and what goes. Pruning only targets toolResult messages. It never modifies your actual user messages or the assistant’s responses.

The system protects the most recent part of your conversation. It looks at the keepLastAssistants setting (which defaults to 3) to establish a cutoff. Any tool results that appear after that cutoff are safe. Pruning also skips any tool results that contain image blocks to avoid breaking visual context.

There are two ways the system trims data:

  1. Soft-trim: This keeps the beginning and end of a large tool result but removes the middle. It adds a note about the original size so the model knows data was removed.
  2. Hard-clear: This replaces the entire tool result with a placeholder like [Old tool result content cleared].

If you don’t see pruning happening, check these common reasons from the documentation:

  • Not enough assistant messages: If the session doesn’t have enough assistant messages to satisfy the keepLastAssistants count, the system skips pruning entirely.
  • Image blocks: If a tool result contains an image, it is protected and will never be trimmed or cleared.
  • Provider compatibility: Pruning currently only runs for Anthropic API calls and Anthropic models on OpenRouter.

If you need help setting this up for your specific workflow, check out the AI Setup Assistant.

OpenClaw

OpenClaw Expert

Still stuck?

If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.