Skip to content

Run OpenClaw Tests and Benchmarks: A Complete Guide

Ever felt the frustration of a broken build after a simple update? We have all been there, and that is why having a solid testing strategy is a lifesaver for any developer.

When you are working with OpenClaw, you can use a full testing kit that includes suites, live tests, and Docker configurations. These OpenClaw tests help you maintain the Gateway and ensure that your contributions are stable before they reach production. You can find the full testing kit details at Testing.

Getting your testing environment right is the first step to a stable contribution. You can use pnpm to run everything from quick unit checks to full coverage reports.

  1. Use pnpm test:force to kill any lingering Gateway process holding the default control port. This command runs the full Vitest suite with an isolated Gateway port so server tests do not collide with a running instance. You should use this when a prior Gateway run left port 18789 occupied.
  2. Run pnpm test:coverage to execute the unit suite with V8 coverage via vitest.unit.config.ts. This is a loaded-file unit coverage gate, not whole-repo all-file coverage. The thresholds are 70% for lines, functions, and statements, and 55% for branches. Because coverage.all is false, the gate measures files loaded by the unit coverage suite instead of treating every split-lane source file as uncovered.
  3. Use pnpm test:coverage:changed to run unit coverage only for files changed since origin/main.
  4. Execute pnpm test:changed to expand changed git paths into scoped Vitest lanes when the diff only touches routable source or test files. Note that config or setup changes still fall back to the native root projects run so wiring edits rerun broadly when needed.
  5. Run pnpm changed:lanes to see the architectural lanes triggered by the diff against origin/main.
  6. Use pnpm check:changed to run the smart changed gate for the diff against origin/main. It runs core work with core test lanes, extension work with extension test lanes, and test-only work with test typecheck or tests only. It also expands public Plugin SDK or plugin-contract changes to extension validation.
  7. Run pnpm test to route explicit file or directory targets through scoped Vitest lanes. Untargeted runs use fixed shard groups and expand to leaf configs for local parallel execution. The extension group always expands to the per-extension shard configs instead of one giant root-project process.
  8. Full and extension shard runs update local timing data in .artifacts/vitest-shard-timings.json. Later runs use those timings to balance slow and fast shards. You can set OPENCLAW_TEST_PROJECTS_TIMINGS=0 to ignore the local timing artifact.
  9. Selected plugin-sdk and commands test files now route through dedicated light lanes that keep only test/setup.ts, leaving runtime-heavy cases on their existing lanes.
  10. Selected plugin-sdk and commands helper source files also map pnpm test:changed to explicit sibling tests in those light lanes. This helps small helper edits avoid rerunning the heavy runtime-backed suites.
  11. The auto-reply feature now splits into three dedicated configs: core, top-level, and reply. This ensures the reply use does not dominate the lighter top-level status, token, or helper tests.
  12. The base Vitest config now defaults to pool: "threads" and isolate: false, with the shared non-isolated runner enabled across the repo configs.
  13. Run pnpm test:channels to execute vitest.channels.config.ts.
  14. Use pnpm test:extensions or pnpm test extensions to run all extension and plugin shards. Heavy channel extensions and OpenAI run as dedicated shards, while other extension groups stay batched. You can use pnpm test extensions/<id> for one bundled plugin lane.
  15. Run pnpm test:perf:imports to enable Vitest import-duration and import-breakdown reporting while still using scoped lane routing for explicit targets.
  16. Use pnpm test:perf:imports:changed for the same import profiling, but only for files changed since origin/main.
  17. Execute pnpm test:perf:changed:bench -- --ref <git-ref> to benchmark the routed changed-mode path against the native root-project run for the same committed git diff.
  18. Run pnpm test:perf:changed:bench -- --worktree to benchmark the current worktree change set without committing first.
  19. Use pnpm test:perf:profile:main to write a CPU profile for the Vitest main thread to .artifacts/vitest-main-profile.
  20. Run pnpm test:perf:profile:runner to write CPU and heap profiles for the unit runner to .artifacts/vitest-runner-profile.
  21. For Gateway integration, you can opt-in via OPENCLAW_TEST_INCLUDE_GATEWAY=1 pnpm test or pnpm test:gateway.
  22. Execute pnpm test:e2e to run Gateway end-to-end smoke tests with multi-instance WS, HTTP, and node pairing. This defaults to threads and isolate: false with adaptive workers in vitest.e2e.config.ts. You can tune this with OPENCLAW_E2E_WORKERS=<n> and set OPENCLAW_E2E_VERBOSE=1 for verbose logs.
  23. Run pnpm test:live to execute provider live tests for minimax or zai. This requires API keys and LIVE=1 or provider-specific flags like *_LIVE_TEST=1 to unskip.
  24. Use pnpm test:docker:openwebui to start a Dockerized OpenClaw and Open WebUI setup. It signs in through Open WebUI, checks /api/models, and runs a real proxied chat through /api/chat/completions. This requires a usable live model key and pulls an external image, so it is not expected to be CI-stable.
  25. Run pnpm test:docker:mcp-channels to start a seeded Gateway container and a client container that spawns openclaw mcp serve. It verifies routed conversation discovery, transcript reads, attachment metadata, live event queue behavior, outbound send routing, and Claude-style channel notifications over the real stdio bridge. The Claude notification assertion reads the raw stdio MCP frames directly so the smoke reflects what the bridge actually emits.

Before you submit your code to GitHub, you should run a series of local checks to ensure everything is in order. These steps help you catch errors early and maintain high code quality.

  1. Run pnpm check:changed.
  2. Run pnpm check.
  3. Run pnpm check:test-types.
  4. Run pnpm build.
  5. Run pnpm test.
  6. Run pnpm check:docs.

If pnpm test flakes on a loaded host, rerun it once before treating it as a regression. You can isolate the issue with pnpm test <path/to/test>. For memory-constrained hosts, use these environment variables:

  • OPENCLAW_VITEST_MAX_WORKERS=1 pnpm test
  • OPENCLAW_VITEST_FS_MODULE_CACHE_PATH=/tmp/openclaw-vitest-cache pnpm test:changed

Understanding the response times of different models helps you optimize your application for speed. You can use the provided script to run benchmarks against your local API keys.

  1. The script is located at scripts/bench-model.ts.
  2. Use the command source ~/.profile && pnpm tsx scripts/bench-model.ts --runs 10 to start the benchmark.
  3. You can set optional environment variables like MINIMAX_API_KEY, MINIMAX_BASE_URL, MINIMAX_MODEL, or ANTHROPIC_API_KEY.
  4. The default prompt used is “Reply with a single word: ok. No punctuation or extra text.”

As of the last run on 2025-12-31 with 20 runs, the minimax median was 1279ms (min 1114, max 2431) and the opus median was 2454ms (min 1224, max 3170).

A fast startup time for the CLI ensures a smooth developer experience. These benchmarking tools allow you to measure and track the performance of various OpenClaw commands.

  1. Use pnpm test:startup:bench for a general benchmark.
  2. Run pnpm test:startup:bench:smoke to write the targeted smoke artifact at .artifacts/cli-startup-bench-smoke.json.
  3. Use pnpm test:startup:bench:save to write the full-suite artifact at .artifacts/cli-startup-bench-all.json using runs=5 and warmup=1.
  4. Run pnpm test:startup:bench:update to refresh the checked-in baseline fixture at test/fixtures/cli-startup-bench.json using runs=5 and warmup=1.
  5. Use pnpm test:startup:bench:check to compare current results against the fixture.
  6. Run pnpm tsx scripts/bench-cli-startup.ts directly for more control.
  7. Use flags like pnpm tsx scripts/bench-cli-startup.ts --runs 12 or --preset real to customize your run.
  8. You can also specify entry points with --entry openclaw.mjs or output results to a JSON file with --json.
  9. Use --cpu-prof-dir .artifacts/cli-cpu to write V8 profiles for each run.

The available presets include:

  • startup: --version, --help, health, health --json, status --json, status
  • real: health, status, status --json, sessions, sessions --json, agents list --json, gateway status, gateway status --json, gateway health --json, config get gateway.port
  • all: both presets

The output provides detailed summaries including sampleCount, average, p50, p95, min/max, exit-code/signal distribution, and max RSS summaries for each command.

Testing the onboarding process in a clean environment ensures that new users can set up OpenClaw without issues. Using Docker allows you to simulate a fresh installation every time.

  1. Run the script scripts/e2e/onboard-docker.sh to start the full cold-start flow in a clean Linux container.
  2. This script drives the interactive wizard via a pseudo-tty and verifies that the config, workspace, and session files are created correctly.
  3. It then starts the Gateway and runs openclaw health to confirm the system is ready.
  4. This ensures the entire onboarding experience is smooth for new developers.
Terminal window
scripts/e2e/onboard-docker.sh

It is important to verify that terminal-based features like QR code generation work across different environments. This test checks compatibility with supported Docker Node.js runtimes.

  1. Run pnpm test:docker:qr to ensure qrcode-terminal loads correctly.
  2. This confirms compatibility with Node.js 24 (default) and Node.js 22.
Terminal window
pnpm test:docker:qr
OpenClaw

OpenClaw Expert

Still stuck?

If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.