OpenClaw Testing — Unit, E2E, and Live Test Suites
OpenClaw Testing
Section titled “OpenClaw Testing”OpenClaw has three test suites, each with increasing realism (and cost). Most days you’ll run the fast unit tests. When debugging real provider issues, you’ll reach for live tests.
Quick Start
Section titled “Quick Start”# Full gate (expected before push)pnpm lint && pnpm build && pnpm test
# With coveragepnpm test:coverage
# E2E suite (gateway networking)pnpm test:e2e
# Live suite (real providers, real costs)pnpm test:liveTest Suites
Section titled “Test Suites”Unit / Integration (Default)
Section titled “Unit / Integration (Default)”pnpm test- Files:
src/**/*.test.ts - Scope: Pure unit tests, in-process integration, deterministic regressions
- Runs in CI: Yes
- Keys required: No
- Speed: Fast ⚡
E2E (Gateway Smoke)
Section titled “E2E (Gateway Smoke)”pnpm test:e2e- Files:
src/**/*.e2e.test.ts - Scope: Multi-instance gateway, WebSocket, HTTP surfaces, node pairing
- Runs in CI: Yes (when enabled)
- Keys required: No
- Speed: Slower
Live (Real Providers)
Section titled “Live (Real Providers)”pnpm test:live- Files:
src/**/*.live.test.ts - Scope: Real API calls to real providers
- Runs in CI: No (not CI-stable by design)
- Keys required: Yes
- Speed: Depends on provider latency
- Cost: Real money / rate limits
When to Use Which
Section titled “When to Use Which”| Scenario | Suite |
|---|---|
| Editing logic/tests | pnpm test |
| Gateway networking changes | Add pnpm test:e2e |
| ”My bot is down” / provider issues | Narrowed pnpm test:live |
Live Testing
Section titled “Live Testing”Live tests have two layers:
Layer 1: Direct Model Completion
Section titled “Layer 1: Direct Model Completion”Test: src/agents/models.profiles.live.test.ts
Tests providers directly without the gateway. Useful for isolating “is the API broken?” from “is my pipeline broken?”
OPENCLAW_LIVE_MODELS="openai/gpt-5.2" pnpm test:live src/agents/models.profiles.live.test.tsLayer 2: Gateway + Agent Smoke
Section titled “Layer 2: Gateway + Agent Smoke”Test: src/gateway/gateway-models.profiles.live.test.ts
Full pipeline: gateway → agent → model → tools.
OPENCLAW_LIVE_GATEWAY_MODELS="openai/gpt-5.2" pnpm test:live src/gateway/gateway-models.profiles.live.test.tsProbes included:
- Read probe: Writes a nonce file, asks agent to read it back
- Exec+read probe: Asks agent to write then read a file
- Image probe: Sends an image, expects model to OCR the content
Narrowing Live Tests
Section titled “Narrowing Live Tests”Always use allowlists to avoid running everything:
# Single model, directOPENCLAW_LIVE_MODELS="openai/gpt-5.2" pnpm test:live src/agents/models.profiles.live.test.ts
# Single model, gatewayOPENCLAW_LIVE_GATEWAY_MODELS="openai/gpt-5.2" pnpm test:live src/gateway/gateway-models.profiles.live.test.ts
# Multiple providersOPENCLAW_LIVE_GATEWAY_MODELS="openai/gpt-5.2,anthropic/claude-opus-4-5,google/gemini-3-flash-preview" pnpm test:liveCredentials
Section titled “Credentials”Live tests find credentials the same way the CLI does:
- Profile store:
~/.openclaw/credentials/ - Config:
~/.openclaw/openclaw.json - Environment variables
# Force profile-only keysOPENCLAW_LIVE_REQUIRE_PROFILE_KEYS=1 pnpm test:liveDocker Runners
Section titled “Docker Runners”Run tests inside Docker for Linux validation:
pnpm test:docker:live-models # Direct modelspnpm test:docker:live-gateway # Gateway + agentpnpm test:docker:onboard # Onboarding wizardpnpm test:docker:gateway-network # Two-container networkingpnpm test:docker:plugins # Plugin loadingDocker Environment Variables
Section titled “Docker Environment Variables”| Variable | Default | Description |
|---|---|---|
OPENCLAW_CONFIG_DIR | ~/.openclaw | Mounted to /home/node/.openclaw |
OPENCLAW_WORKSPACE_DIR | ~/.openclaw/workspace | Mounted to /home/node/.openclaw/workspace |
OPENCLAW_PROFILE_FILE | ~/.profile | Sourced before running tests |
OPENCLAW_LIVE_GATEWAY_MODELS | — | Narrow model selection |
OPENCLAW_LIVE_MODELS | — | Narrow model selection (direct) |
OPENCLAW_LIVE_REQUIRE_PROFILE_KEYS | 0 | Force profile-only credentials |
Docker Scripts
Section titled “Docker Scripts”| Command | Script | Purpose |
|---|---|---|
test:docker:live-models | scripts/test-live-models-docker.sh | Direct model completion |
test:docker:live-gateway | scripts/test-live-gateway-models-docker.sh | Gateway + dev agent |
test:docker:onboard | scripts/e2e/onboard-docker.sh | TTY onboarding wizard |
test:docker:gateway-network | scripts/e2e/gateway-network-docker.sh | Two-container WS auth |
test:docker:plugins | scripts/e2e/plugins-docker.sh | Custom extension loading |
Adding Regressions
Section titled “Adding Regressions”When you fix a provider/model issue:
- CI-safe first: Mock/stub the provider if possible
- Live-only if necessary: Keep it narrow and env-gated
- Target the right layer:
- Provider bug →
models.profiles.live.test.ts - Pipeline bug →
gateway-models.profiles.live.test.ts
- Provider bug →
Recommended Live Recipes
Section titled “Recommended Live Recipes”Narrow, explicit allowlists are fastest and least flaky:
# Single model, direct (no gateway)OPENCLAW_LIVE_MODELS="openai/gpt-5.2" pnpm test:live src/agents/models.profiles.live.test.ts
# Single model, gateway smokeOPENCLAW_LIVE_GATEWAY_MODELS="openai/gpt-5.2" pnpm test:live src/gateway/gateway-models.profiles.live.test.ts
# Tool calling across several providersOPENCLAW_LIVE_GATEWAY_MODELS="openai/gpt-5.2,anthropic/claude-opus-4-5,google/gemini-3-flash-preview,zai/glm-4.7,minimax/minimax-m2.1" pnpm test:live src/gateway/gateway-models.profiles.live.test.ts
# Google focus (Gemini API key + Antigravity)OPENCLAW_LIVE_GATEWAY_MODELS="google/gemini-3-flash-preview" pnpm test:live src/gateway/gateway-models.profiles.live.test.tsOPENCLAW_LIVE_GATEWAY_MODELS="google-antigravity/claude-opus-4-5-thinking" pnpm test:live src/gateway/gateway-models.profiles.live.test.tsAnthropic Setup-Token Smoke
Section titled “Anthropic Setup-Token Smoke”Test: src/agents/anthropic.setup-token.live.test.ts
Verify Claude Code CLI setup-token can complete an Anthropic prompt.
# EnableOPENCLAW_LIVE_SETUP_TOKEN=1 pnpm test:live src/agents/anthropic.setup-token.live.test.ts
# With profileOPENCLAW_LIVE_SETUP_TOKEN_PROFILE=anthropic:setup-token-test pnpm test:live
# With raw tokenOPENCLAW_LIVE_SETUP_TOKEN_VALUE=sk-ant-oat01-... pnpm test:liveSetup:
openclaw models auth paste-token --provider anthropic --profile-id anthropic:setup-token-testCLI Backend Smoke
Section titled “CLI Backend Smoke”Test: src/gateway/gateway-cli-backend.live.test.ts
Validate the Gateway + agent pipeline using a local CLI backend (like Claude Code CLI).
# BasicOPENCLAW_LIVE_CLI_BACKEND=1 pnpm test:live src/gateway/gateway-cli-backend.live.test.ts
# With model overrideOPENCLAW_LIVE_CLI_BACKEND=1 \ OPENCLAW_LIVE_CLI_BACKEND_MODEL="claude-cli/claude-sonnet-4-5" \ pnpm test:live src/gateway/gateway-cli-backend.live.test.tsEnvironment Variables:
| Variable | Default | Description |
|---|---|---|
OPENCLAW_LIVE_CLI_BACKEND_MODEL | claude-cli/claude-sonnet-4-5 | Model to use |
OPENCLAW_LIVE_CLI_BACKEND_COMMAND | claude | CLI command path |
OPENCLAW_LIVE_CLI_BACKEND_ARGS | ["-p","--output-format","json","--dangerously-skip-permissions"] | CLI args |
OPENCLAW_LIVE_CLI_BACKEND_IMAGE_PROBE | 0 | Enable image attachment test |
OPENCLAW_LIVE_CLI_BACKEND_RESUME_PROBE | 0 | Enable multi-turn test |
Model Matrix
Section titled “Model Matrix”Modern Smoke Set (Recommended)
Section titled “Modern Smoke Set (Recommended)”| Provider | Model |
|---|---|
| OpenAI | openai/gpt-5.2 |
| OpenAI Codex | openai-codex/gpt-5.2 |
| Anthropic | anthropic/claude-opus-4-5 |
| Google (API) | google/gemini-3-flash-preview |
| Google (Antigravity) | google-antigravity/claude-opus-4-5-thinking |
| Z.AI | zai/glm-4.7 |
| MiniMax | minimax/minimax-m2.1 |
Run All Modern Models
Section titled “Run All Modern Models”OPENCLAW_LIVE_GATEWAY_MODELS="openai/gpt-5.2,openai-codex/gpt-5.2,anthropic/claude-opus-4-5,google/gemini-3-flash-preview,zai/glm-4.7,minimax/minimax-m2.1" pnpm test:live src/gateway/gateway-models.profiles.live.test.tsCheck Available Models
Section titled “Check Available Models”openclaw models listopenclaw models list --jsonDeepgram Live (Audio Transcription)
Section titled “Deepgram Live (Audio Transcription)”Test: src/media-understanding/providers/deepgram/audio.live.test.ts
DEEPGRAM_API_KEY=... DEEPGRAM_LIVE_TEST=1 pnpm test:live src/media-understanding/providers/deepgram/audio.live.test.tsDocs Sanity
Section titled “Docs Sanity”Run docs checks after doc edits:
pnpm docs:listOffline Regression (CI-Safe)
Section titled “Offline Regression (CI-Safe)”These run in CI without real providers:
| Test | Description |
|---|---|
gateway.tool-calling.mock-openai.test.ts | Gateway tool calling with mock OpenAI |
gateway.wizard.e2e.test.ts | Wizard WebSocket + config writes |
Agent Reliability Evals
Section titled “Agent Reliability Evals”Current coverage:
- Mock tool-calling through real gateway + agent loop
- End-to-end wizard flows with session wiring
What’s planned:
- Decisioning: Does agent pick the right skill?
- Compliance: Does agent follow skill instructions?
- Workflow contracts: Multi-turn tool order assertions
Still stuck? Our AI Setup Assistant can help with test issues.
What’s Next?
Section titled “What’s Next?”- Logging → — Log files and console output
- Debugging → — Watch mode and raw streams
- Contributing → — Development guidelines
Need help? Join the OpenClaw Discord or check the GitHub Issues.
OpenClaw Expert
Still stuck?
If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.