Run OpenClaw Tests and Benchmarks: A Complete Guide
Ever felt the frustration of a broken build after a simple update? We have all been there, and that is why having a solid testing strategy is a lifesaver for any developer.
When you are working with OpenClaw, you can use a full testing kit that includes suites, live tests, and Docker configurations. These OpenClaw tests help you maintain the Gateway and ensure that your contributions are stable before they reach production. You can find the full testing kit details at Testing.
Run OpenClaw Test Suites
Section titled “Run OpenClaw Test Suites”Getting your testing environment right is the first step to a stable contribution. You can use pnpm to run everything from quick unit checks to full coverage reports.
- Use
pnpm test:forceto kill any lingering Gateway process holding the default control port. This command runs the full Vitest suite with an isolated Gateway port so server tests do not collide with a running instance. You should use this when a prior Gateway run left port 18789 occupied. - Run
pnpm test:coverageto execute the unit suite with V8 coverage viavitest.unit.config.ts. This is a loaded-file unit coverage gate, not whole-repo all-file coverage. The thresholds are 70% for lines, functions, and statements, and 55% for branches. Becausecoverage.allis false, the gate measures files loaded by the unit coverage suite instead of treating every split-lane source file as uncovered. - Use
pnpm test:coverage:changedto run unit coverage only for files changed sinceorigin/main. - Execute
pnpm test:changedto expand changed git paths into scoped Vitest lanes when the diff only touches routable source or test files. Note that config or setup changes still fall back to the native root projects run so wiring edits rerun broadly when needed. - Run
pnpm changed:lanesto see the architectural lanes triggered by the diff againstorigin/main. - Use
pnpm check:changedto run the smart changed gate for the diff againstorigin/main. It runs core work with core test lanes, extension work with extension test lanes, and test-only work with test typecheck or tests only. It also expands public Plugin SDK or plugin-contract changes to extension validation. - Run
pnpm testto route explicit file or directory targets through scoped Vitest lanes. Untargeted runs use fixed shard groups and expand to leaf configs for local parallel execution. The extension group always expands to the per-extension shard configs instead of one giant root-project process. - Full and extension shard runs update local timing data in
.artifacts/vitest-shard-timings.json. Later runs use those timings to balance slow and fast shards. You can setOPENCLAW_TEST_PROJECTS_TIMINGS=0to ignore the local timing artifact. - Selected
plugin-sdkandcommandstest files now route through dedicated light lanes that keep onlytest/setup.ts, leaving runtime-heavy cases on their existing lanes. - Selected
plugin-sdkandcommandshelper source files also mappnpm test:changedto explicit sibling tests in those light lanes. This helps small helper edits avoid rerunning the heavy runtime-backed suites. - The
auto-replyfeature now splits into three dedicated configs:core,top-level, andreply. This ensures the reply use does not dominate the lighter top-level status, token, or helper tests. - The base Vitest config now defaults to
pool: "threads"andisolate: false, with the shared non-isolated runner enabled across the repo configs. - Run
pnpm test:channelsto executevitest.channels.config.ts. - Use
pnpm test:extensionsorpnpm test extensionsto run all extension and plugin shards. Heavy channel extensions and OpenAI run as dedicated shards, while other extension groups stay batched. You can usepnpm test extensions/<id>for one bundled plugin lane. - Run
pnpm test:perf:importsto enable Vitest import-duration and import-breakdown reporting while still using scoped lane routing for explicit targets. - Use
pnpm test:perf:imports:changedfor the same import profiling, but only for files changed sinceorigin/main. - Execute
pnpm test:perf:changed:bench -- --ref <git-ref>to benchmark the routed changed-mode path against the native root-project run for the same committed git diff. - Run
pnpm test:perf:changed:bench -- --worktreeto benchmark the current worktree change set without committing first. - Use
pnpm test:perf:profile:mainto write a CPU profile for the Vitest main thread to.artifacts/vitest-main-profile. - Run
pnpm test:perf:profile:runnerto write CPU and heap profiles for the unit runner to.artifacts/vitest-runner-profile. - For Gateway integration, you can opt-in via
OPENCLAW_TEST_INCLUDE_GATEWAY=1 pnpm testorpnpm test:gateway. - Execute
pnpm test:e2eto run Gateway end-to-end smoke tests with multi-instance WS, HTTP, and node pairing. This defaults tothreadsandisolate: falsewith adaptive workers invitest.e2e.config.ts. You can tune this withOPENCLAW_E2E_WORKERS=<n>and setOPENCLAW_E2E_VERBOSE=1for verbose logs. - Run
pnpm test:liveto execute provider live tests for minimax or zai. This requires API keys andLIVE=1or provider-specific flags like*_LIVE_TEST=1to unskip. - Use
pnpm test:docker:openwebuito start a Dockerized OpenClaw and Open WebUI setup. It signs in through Open WebUI, checks/api/models, and runs a real proxied chat through/api/chat/completions. This requires a usable live model key and pulls an external image, so it is not expected to be CI-stable. - Run
pnpm test:docker:mcp-channelsto start a seeded Gateway container and a client container that spawnsopenclaw mcp serve. It verifies routed conversation discovery, transcript reads, attachment metadata, live event queue behavior, outbound send routing, and Claude-style channel notifications over the real stdio bridge. The Claude notification assertion reads the raw stdio MCP frames directly so the smoke reflects what the bridge actually emits.
Perform Local PR Gate Checks
Section titled “Perform Local PR Gate Checks”Before you submit your code to GitHub, you should run a series of local checks to ensure everything is in order. These steps help you catch errors early and maintain high code quality.
- Run
pnpm check:changed. - Run
pnpm check. - Run
pnpm check:test-types. - Run
pnpm build. - Run
pnpm test. - Run
pnpm check:docs.
If pnpm test flakes on a loaded host, rerun it once before treating it as a regression. You can isolate the issue with pnpm test <path/to/test>. For memory-constrained hosts, use these environment variables:
OPENCLAW_VITEST_MAX_WORKERS=1 pnpm testOPENCLAW_VITEST_FS_MODULE_CACHE_PATH=/tmp/openclaw-vitest-cache pnpm test:changed
Benchmark Model Latency with Local Keys
Section titled “Benchmark Model Latency with Local Keys”Understanding the response times of different models helps you optimize your application for speed. You can use the provided script to run benchmarks against your local API keys.
- The script is located at
scripts/bench-model.ts. - Use the command
source ~/.profile && pnpm tsx scripts/bench-model.ts --runs 10to start the benchmark. - You can set optional environment variables like
MINIMAX_API_KEY,MINIMAX_BASE_URL,MINIMAX_MODEL, orANTHROPIC_API_KEY. - The default prompt used is “Reply with a single word: ok. No punctuation or extra text.”
As of the last run on 2025-12-31 with 20 runs, the minimax median was 1279ms (min 1114, max 2431) and the opus median was 2454ms (min 1224, max 3170).
Analyze CLI Startup Performance
Section titled “Analyze CLI Startup Performance”A fast startup time for the CLI ensures a smooth developer experience. These benchmarking tools allow you to measure and track the performance of various OpenClaw commands.
- Use
pnpm test:startup:benchfor a general benchmark. - Run
pnpm test:startup:bench:smoketo write the targeted smoke artifact at.artifacts/cli-startup-bench-smoke.json. - Use
pnpm test:startup:bench:saveto write the full-suite artifact at.artifacts/cli-startup-bench-all.jsonusingruns=5andwarmup=1. - Run
pnpm test:startup:bench:updateto refresh the checked-in baseline fixture attest/fixtures/cli-startup-bench.jsonusingruns=5andwarmup=1. - Use
pnpm test:startup:bench:checkto compare current results against the fixture. - Run
pnpm tsx scripts/bench-cli-startup.tsdirectly for more control. - Use flags like
pnpm tsx scripts/bench-cli-startup.ts --runs 12or--preset realto customize your run. - You can also specify entry points with
--entry openclaw.mjsor output results to a JSON file with--json. - Use
--cpu-prof-dir .artifacts/cli-cputo write V8 profiles for each run.
The available presets include:
startup:--version,--help,health,health --json,status --json,statusreal:health,status,status --json,sessions,sessions --json,agents list --json,gateway status,gateway status --json,gateway health --json,config get gateway.portall: both presets
The output provides detailed summaries including sampleCount, average, p50, p95, min/max, exit-code/signal distribution, and max RSS summaries for each command.
Run Onboarding E2E Tests in Docker
Section titled “Run Onboarding E2E Tests in Docker”Testing the onboarding process in a clean environment ensures that new users can set up OpenClaw without issues. Using Docker allows you to simulate a fresh installation every time.
- Run the script
scripts/e2e/onboard-docker.shto start the full cold-start flow in a clean Linux container. - This script drives the interactive wizard via a pseudo-tty and verifies that the config, workspace, and session files are created correctly.
- It then starts the Gateway and runs
openclaw healthto confirm the system is ready. - This ensures the entire onboarding experience is smooth for new developers.
scripts/e2e/onboard-docker.shExecute QR Import Smoke Tests
Section titled “Execute QR Import Smoke Tests”It is important to verify that terminal-based features like QR code generation work across different environments. This test checks compatibility with supported Docker Node.js runtimes.
- Run
pnpm test:docker:qrto ensureqrcode-terminalloads correctly. - This confirms compatibility with Node.js 24 (default) and Node.js 22.
pnpm test:docker:qrNext Steps
Section titled “Next Steps”OpenClaw Expert
Still stuck?
If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.