Master OpenClaw Testing: Run Unit, E2E, and Live Suites
Get started with OpenClaw testing
Section titled “Get started with OpenClaw testing”OpenClaw testing is designed to be efficient and reliable, whether you are doing a quick check or a full validation of your changes. You can use these commands to maintain high code quality during your daily development workflow and ensure your Gateway is always performing as expected.
- Run the full gate, which is expected before you push any code:
pnpm build && pnpm check && pnpm check:test-types && pnpm test - Execute a faster local full-suite run if you are working on a machine with plenty of resources:
pnpm test:max - Use the direct Vitest watch loop for active development:
pnpm test:watch - Target specific files directly, which now also routes extension and channel paths correctly:
pnpm test extensions/discord/src/monitor/message-handler.preflight.test.ts - Iterate on a single failure by using targeted runs first before running the entire suite.
- Start the Docker-backed QA site for manual testing:
pnpm qa:lab:up - Run the Linux VM-backed QA lane using Multipass:
pnpm openclaw qa suite --runner multipass --scenario channel-chat-baseline
When you touch tests or want extra confidence in your changes:
- Run the coverage gate to check your code coverage:
pnpm test:coverage - Execute the E2E suite for integration testing:
pnpm test:e2e
When you are debugging real providers or models, which requires real credentials:
- Run the live suite for models and Gateway tool or image probes:
pnpm test:live - Target a single live file quietly to focus on one issue:
pnpm test:live -- src/agents/models.profiles.live.test.ts
Tip: When you only need to fix one failing case, you should prefer narrowing down the live tests using the allowlist environment variables described below.
Run QA-specific scenarios in OpenClaw
Section titled “Run QA-specific scenarios in OpenClaw”These commands provide a higher level of realism for your QA-lab work by simulating actual environments and transport layers. They sit alongside the main test suites to ensure your Gateway handles real-world scenarios correctly across different platforms.
pnpm openclaw qa suite: This runs repo-backed QA scenarios directly on your host. It runs multiple selected scenarios in parallel by default with isolated Gateway workers. Theqa-channeldefaults to a concurrency of 4, which is bounded by the selected scenario count. You can tune the worker count with--concurrency <count>or use--concurrency 1for the older serial lane. It exits with a non-zero code if any scenario fails, unless you use--allow-failureswhen you want artifacts without a failing exit code. It supports provider modes likelive-frontier,mock-openai, andaimock. Theaimockmode starts a local AIMock-backed provider server for experimental fixture and protocol-mock coverage.pnpm openclaw qa suite --runner multipass: This runs the same QA suite inside a disposable Multipass Linux VM. It keeps the same scenario-selection behavior as the host-basedqa suiteand reuses the same provider/model selection flags. Live runs forward the supported QA auth inputs that are practical for the guest, such as environment-based provider keys, the QA live provider config path, andCODEX_HOMEwhen present. Output directories must stay under the repo root so the guest can write back through the mounted workspace. It writes the normal QA report, summary, and Multipass logs under.artifacts/qa-e2e/....pnpm qa:lab:up: This starts the Docker-backed QA site for operator-style QA work.pnpm openclaw qa aimock: This starts only the local AIMock provider server for direct protocol smoke testing.pnpm openclaw qa matrix: This runs the Matrix live QA lane against a disposable Docker-backed Tuwunel homeserver. This QA host is currently for repo and development use only, as packaged OpenClaw installs do not shipqa-lab. It provisions three temporary Matrix users (driver,sut,observer) and one private room, then starts a QA Gateway child with the real Matrix plugin as the SUT transport. It uses the pinned stable Tuwunel imageghcr.io/matrix-construct/tuwunel:v1.5.1by default, which you can override withOPENCLAW_QA_MATRIX_TUWUNEL_IMAGE. It writes a Matrix QA report, summary, observed-events artifact, and combined logs under.artifacts/qa-e2e/....pnpm openclaw qa telegram: This runs the Telegram live QA lane against a real private group using the driver and SUT bot tokens from your environment. It requiresOPENCLAW_QA_TELEGRAM_GROUP_ID,OPENCLAW_QA_TELEGRAM_DRIVER_BOT_TOKEN, andOPENCLAW_QA_TELEGRAM_SUT_BOT_TOKEN. It supports--credential-source convexfor shared pooled credentials. It requires two distinct bots in the same private group, with the SUT bot exposing a Telegram username. For stable bot-to-bot observation, you should enable Bot-to-Bot Communication Mode in@BotFatherfor both bots. It writes a Telegram QA report, summary, and observed-messages artifact under.artifacts/qa-e2e/....
Live transport lanes share one standard contract so new transports do not drift:
qa-channel remains the broad synthetic QA suite and is not part of the live transport coverage matrix.
| Lane | Canary | Mention gating | Allowlist block | Top-level reply | Restart resume | Thread follow-up | Thread isolation | Reaction observation | Help command |
|---|---|---|---|---|---|---|---|---|---|
| Matrix | x | x | x | x | x | x | x | x | |
| Telegram | x | x |
Shared Telegram credentials via Convex (v1)
Section titled “Shared Telegram credentials via Convex (v1)”When you enable --credential-source convex (or set OPENCLAW_QA_CREDENTIAL_SOURCE=convex) for openclaw qa telegram, the QA lab acquires an exclusive lease from a Convex-backed pool. It heartbeats that lease while the lane is running and releases it upon shutdown.
Reference Convex project scaffold:
qa/convex-credential-broker/
Required environment variables:
OPENCLAW_QA_CONVEX_SITE_URL(for examplehttps://your-deployment.convex.site)- One secret for the selected role:
OPENCLAW_QA_CONVEX_SECRET_MAINTAINERformaintainerorOPENCLAW_QA_CONVEX_SECRET_CIforci. - Credential role selection: Use the CLI flag
--credential-role maintainer|cior the environment defaultOPENCLAW_QA_CREDENTIAL_ROLE(defaults tociin CI,maintainerotherwise).
Optional environment variables:
OPENCLAW_QA_CREDENTIAL_LEASE_TTL_MS(default1200000)OPENCLAW_QA_CREDENTIAL_HEARTBEAT_INTERVAL_MS(default30000)OPENCLAW_QA_CREDENTIAL_ACQUIRE_TIMEOUT_MS(default90000)OPENCLAW_QA_CREDENTIAL_HTTP_TIMEOUT_MS(default15000)OPENCLAW_QA_CONVEX_ENDPOINT_PREFIX(default/qa-credentials/v1)OPENCLAW_QA_CREDENTIAL_OWNER_ID(optional trace id)OPENCLAW_QA_ALLOW_INSECURE_HTTP=1allows loopbackhttp://Convex URLs for local-only development.
Maintainer admin commands for pool management require OPENCLAW_QA_CONVEX_SECRET_MAINTAINER specifically. You can use these CLI helpers:
pnpm openclaw qa credentials add --kind telegram --payload-file qa/telegram-credential.jsonpnpm openclaw qa credentials list --kind telegrampnpm openclaw qa credentials remove --credential-id <credential-id>Use --json for machine-readable output in scripts and CI utilities. The default endpoint contract (OPENCLAW_QA_CONVEX_SITE_URL + /qa-credentials/v1) includes:
POST /acquire: Request includes{ kind, ownerId, actorRole, leaseTtlMs, heartbeatIntervalMs }. Success returns{ status: "ok", credentialId, leaseToken, payload, leaseTtlMs?, heartbeatIntervalMs? }.POST /heartbeat: Request includes{ kind, ownerId, actorRole, credentialId, leaseToken, leaseTtlMs }. Success returns{ status: "ok" }.POST /release: Request includes{ kind, ownerId, actorRole, credentialId, leaseToken }. Success returns{ status: "ok" }.POST /admin/add: Request includes{ kind, actorId, payload, note?, status? }(maintainer secret only).POST /admin/remove: Request includes{ credentialId, actorId }(maintainer secret only).POST /admin/list: Request includes{ kind?, status?, includePayload?, limit? }(maintainer secret only).
The JSON payload shape for the Telegram kind consists of:
{ groupId: string, driverToken: string, sutToken: string }groupIdmust be a numeric Telegram chat id string.admin/addvalidates this shape forkind: "telegram"and rejects malformed payloads.
Adding a channel to QA
Section titled “Adding a channel to QA”Adding a channel to the markdown QA system requires exactly two things: a transport adapter for the channel and a scenario pack that exercises the channel contract. You should not add a new top-level QA command root when the shared qa-lab host can own the flow.
The qa-lab owns the shared host mechanics:
- The
openclaw qacommand root. - Suite startup and teardown.
- Worker concurrency and artifact writing.
- Report generation and scenario execution.
- Compatibility aliases for older
qa-channelscenarios.
Runner plugins own the transport contract:
- How
openclaw qa <runner>is mounted beneath the sharedqaroot. - How the Gateway is configured for that transport.
- How readiness is checked and how inbound events are injected.
- How outbound messages are observed and how transcripts are exposed.
- How transport-backed actions are executed and how cleanup is handled.
The minimum adoption bar for a new channel includes:
- Keep
qa-labas the owner of the sharedqaroot. - Implement the transport runner on the shared
qa-labhost seam. - Keep transport-specific mechanics inside the runner plugin or channel use.
- Mount the runner as
openclaw qa <runner>instead of registering a competing root command. Runner plugins should declareqaRunnersinopenclaw.plugin.jsonand export a matchingqaRunnerCliRegistrationsarray fromruntime-api.ts. - Author or adapt markdown scenarios under the themed
qa/scenarios/directories. - Use the generic scenario helpers for new scenarios.
- Keep existing compatibility aliases working unless the repo is doing an intentional migration.
Preferred generic helper names for new scenarios are:
waitForTransportReadywaitForChannelReadyinjectInboundMessageinjectOutboundMessagewaitForTransportOutboundMessagewaitForChannelOutboundMessagewaitForNoTransportOutboundgetTransportSnapshotreadTransportMessagereadTransportTranscriptformatTransportTranscriptresetTransport
Compatibility aliases remain available for existing scenarios, including:
waitForQaChannelReadywaitForOutboundMessagewaitForNoOutboundformatConversationTranscriptresetBus
Understand OpenClaw test suites
Section titled “Understand OpenClaw test suites”OpenClaw uses a tiered testing strategy where each suite offers increasing realism at the cost of speed and complexity. You should think of these suites as a progression that helps you catch everything from simple logic errors to complex provider integration issues.
Unit / integration (default)
Section titled “Unit / integration (default)”- Command:
pnpm test - Config: Ten sequential shard runs (
vitest.full-*.config.ts) over the existing scoped Vitest projects. - Files: Core/unit inventories under
src/**/*.test.ts,packages/**/*.test.ts,test/**/*.test.ts, and the whitelisteduinode tests. - Scope: Pure unit tests, in-process integration tests (for Gateway auth, routing, tooling, parsing, and config), and deterministic regressions for known bugs.
- Expectations: These run in CI, require no real keys, and should be fast and stable.
Technical notes for the unit suite:
- Sharding: Untargeted
pnpm testnow runs eleven smaller shard configs (likecore-unit-src,agentic,extensions) instead of one giant process. This cuts peak RSS on loaded machines and avoids resource starvation. - Watch Mode:
pnpm test --watchstill uses the native rootvitest.config.tsproject graph because a multi-shard watch loop is not practical. - Targeted Runs:
pnpm test,pnpm test:watch, andpnpm test:perf:importsroute explicit file targets through scoped lanes first to avoid the full root project startup tax. - Changed Files:
pnpm test:changedexpands changed git paths into the same scoped lanes when the diff only touches routable source/test files.pnpm check:changedis the normal smart local gate that classifies the diff into core, extensions, apps, and more. - Fast Lanes: Import-light unit tests from agents, commands, and
plugin-sdkroute through theunit-fastlane, which skipstest/setup-openclaw-runtime.ts. - Auto-reply: This now has three dedicated buckets to keep the heaviest reply use work off the cheaper status and token tests.
- Embedded Runner: When you change message-tool discovery inputs, you must keep both levels of coverage. Focused helper regressions and embedded runner integration suites (like
src/agents/pi-embedded-runner/compact.hooks.test.ts) verify that behavior flows through the real paths. - Pool and Isolation: Base Vitest config now defaults to
threadsand usesisolate: falseacross the root projects, e2e, and live configs. The sharedscripts/run-vitest.mjslauncher adds--no-maglevfor Vitest child Node.js processes by default to reduce V8 compile churn. - Local Iteration:
pnpm changed:lanesshows which architectural lanes a diff triggers. Local worker auto-scaling is conservative and backs off when the host load average is high. - Performance Debugging:
pnpm test:perf:importsenables Vitest import-duration reporting.pnpm test:perf:changed:bench -- --ref <git-ref>compares routedtest:changedagainst the native root-project path.pnpm test:perf:profile:mainwrites a main-thread CPU profile for Vitest and Vite startup overhead.
E2E (gateway smoke)
Section titled “E2E (gateway smoke)”- Command:
pnpm test:e2e - Config:
vitest.e2e.config.ts - Files:
src/**/*.e2e.test.ts,test/**/*.e2e.test.ts - Scope: Multi-instance Gateway end-to-end behavior, WebSocket and HTTP surfaces, node pairing, and heavier networking.
- Expectations: Runs in CI (when enabled), requires no real keys, but has more moving parts than unit tests.
- Overrides: Use
OPENCLAW_E2E_WORKERS=<n>to force worker count (capped at 16) orOPENCLAW_E2E_VERBOSE=1to re-enable verbose console output.
E2E: OpenShell backend smoke
Section titled “E2E: OpenShell backend smoke”- Command:
pnpm test:e2e:openshell - File:
test/openshell-sandbox.e2e.test.ts - Scope: Starts an isolated OpenShell Gateway on the host via Docker, creates a sandbox from a temporary local Dockerfile, and exercises the OpenShell backend over real SSH.
- Expectations: This is opt-in only and requires a local
openshellCLI plus a working Docker daemon. - Overrides: Set
OPENCLAW_E2E_OPENSHELL=1to enable the test orOPENCLAW_E2E_OPENSHELL_COMMAND=/path/to/openshellto point at a specific binary.
Live (real providers + real models)
Section titled “Live (real providers + real models)”- Command:
pnpm test:live - Config:
vitest.live.config.ts - Files:
src/**/*.live.test.ts - Scope: Verifies if a provider or model actually works today with real credentials. It catches provider format changes, tool-calling quirks, auth issues, and rate limit behavior.
- Expectations: Not CI-stable by design due to real network and provider policies. It costs money and uses rate limits, so you should prefer running narrowed subsets.
- Key Management: Live runs source
~/.profileto pick up missing API keys. By default, they isolateHOMEand copy config into a temp test home. SetOPENCLAW_LIVE_USE_REAL_HOME=1only if you intentionally need to use your real home directory. - Quiet Mode:
pnpm test:livedefaults to a quieter mode, suppressing extra notices and Gateway bootstrap logs. SetOPENCLAW_LIVE_TEST_QUIET=0for full logs. - Key Rotation: Set
*_API_KEYSwith comma/semicolon format or*_API_KEY_1,*_API_KEY_2for provider-specific rotation. Tests will retry on rate limit responses. - Progress Output: Live suites emit progress lines to stderr so long provider calls are visibly active. You can tune direct-model heartbeats with
OPENCLAW_LIVE_HEARTBEAT_MSand Gateway probes withOPENCLAW_LIVE_GATEWAY_HEARTBEAT_MS.
Choose the right OpenClaw test suite
Section titled “Choose the right OpenClaw test suite”Deciding which test to run depends on the scope of your changes and the level of validation you need. Use this simple decision guide to pick the most efficient path for your current task.
- If you are editing logic or standard tests: Run
pnpm test. If you have changed a significant amount of code, you should also runpnpm test:coverage. - If you are touching Gateway networking, the WebSocket protocol, or node pairing: You should add
pnpm test:e2eto your validation process. - If you are debugging “my bot is down” issues, provider-specific failures, or complex tool calling: Run a narrowed subset of
pnpm test:liveto test against real models.
Test OpenClaw Android Node Capabilities
Section titled “Test OpenClaw Android Node Capabilities”You can use the OpenClaw Android live test suite to verify that your connected mobile nodes are handling commands correctly. This ensures every advertised capability actually works as expected in a real-world environment.
- Locate the test file at
src/gateway/android-node.capabilities.live.test.ts. - Run the integration script using the command pnpm android:test:integration.
- The primary goal is to invoke every command currently advertised by a connected Android node and assert command contract behavior.
- Note that the scope involves preconditioned or manual setup, as the suite does not install, run, or pair the app automatically.
- The suite performs command-by-command Gateway
node.invokevalidation for your selected Android node.
To ensure the tests run correctly, you need to complete the following pre-setup:
- Confirm the Android app is already connected and paired to the Gateway.
- Keep the app in the foreground during the testing process.
- Grant all necessary permissions and capture consent for the capabilities you expect to pass.
You can also use optional target overrides to customize the test:
- Set
OPENCLAW_ANDROID_NODE_IDorOPENCLAW_ANDROID_NODE_NAMEto target a specific device. - Use
OPENCLAW_ANDROID_GATEWAY_URL,OPENCLAW_ANDROID_GATEWAY_TOKEN, orOPENCLAW_ANDROID_GATEWAY_PASSWORDto define connection parameters. - Refer to the Android App documentation for full setup details.
Run OpenClaw Model Smoke Tests with Profile Keys
Section titled “Run OpenClaw Model Smoke Tests with Profile Keys”Testing models in OpenClaw is split into two distinct layers to help you quickly identify if a failure is due to the provider’s API or the internal pipeline. This isolation makes debugging much easier when you are dealing with multiple LLM providers.
Layer 1: Direct model completion (no gateway)
Section titled “Layer 1: Direct model completion (no gateway)”- Use the test at
src/agents/models.profiles.live.test.tsto validate direct model access. - The goal is to enumerate discovered models and use
getApiKeyForModelto select models you have credentials for. - This layer runs a small completion per model and includes targeted regressions where needed.
- Enable this by running pnpm test:live or setting
OPENCLAW_LIVE_TEST=1if you are invoking Vitest directly. - You must set
OPENCLAW_LIVE_MODELS=modern(orall) to actually run this suite; otherwise, it skips to keep pnpm test:live focused on Gateway smoke. - Select models using
OPENCLAW_LIVE_MODELS=modernfor the modern allowlist, which includes Opus/Sonnet 4.6+, GPT-5.x + Codex, Gemini 3, GLM 4.7, MiniMax M2.7, and Grok 4. - Use a comma-separated allowlist like
OPENCLAW_LIVE_MODELS="openai/gpt-5.4,anthropic/claude-opus-4-6,..."for specific targets. - Modern sweeps default to a curated high-signal cap, but you can set
OPENCLAW_LIVE_MAX_MODELS=0for an exhaustive sweep or a positive number for a smaller cap. - Select specific providers using
OPENCLAW_LIVE_PROVIDERS="google,google-antigravity,google-gemini-cli". - By default, keys come from the profile store and env fallbacks, but you can set
OPENCLAW_LIVE_REQUIRE_PROFILE_KEYS=1to enforce profile store only. - This layer exists to separate “provider API is broken or key is invalid” from “Gateway agent pipeline is broken” and contains small regressions like OpenAI or Codex reasoning replays.
Layer 2: Gateway + dev agent smoke
Section titled “Layer 2: Gateway + dev agent smoke”- This test is located at
src/gateway/gateway-models.profiles.live.test.tsand covers what “@openclaw” actually does. - The goal is to spin up an in-process Gateway and create or patch an
agent:dev:*session with a model override per run. - It iterates through models-with-keys to assert a meaningful response without tools, a real tool invocation, and optional extra tool probes.
- The
readprobe involves the test writing a nonce file in the workspace and asking the agent to read it and echo the nonce back. - The
exec+readprobe asks the agent to useexecto write a nonce into a temp file and then read it back. - The image probe attaches a generated PNG containing a cat and randomized code, expecting the model to return “cat” plus the code.
- Implementation details are found in
src/gateway/gateway-models.profiles.live.test.tsandsrc/gateway/live-image-probe.ts. - Enable this layer using pnpm test:live.
- Select models via
OPENCLAW_LIVE_GATEWAY_MODELS=allor a specific comma-separated list of “provider/model” strings. - Use
OPENCLAW_LIVE_GATEWAY_PROVIDERSto narrow the providers and avoid testing everything through OpenRouter. - Tool and image probes are always active in this test; the image probe runs when the model advertises image input support.
- The high-level flow involves generating a tiny PNG, sending it via
attachments, parsing it intoimages[]in the Gateway, and forwarding a multimodal message to the model. - The assertion allows for minor OCR mistakes when the model replies with the code.
If you want to see which models you can test on your machine and their exact IDs, run these commands:
openclaw models listopenclaw models list --jsonValidate OpenClaw CLI Backends for Claude and Codex
Section titled “Validate OpenClaw CLI Backends for Claude and Codex”You can validate the OpenClaw Gateway and agent pipeline using local CLI backends without touching your default configuration. This is perfect for testing tools like Claude, Codex, or Gemini in a controlled environment.
- Use the test at
src/gateway/gateway-cli-backend.live.test.tsto validate the pipeline with a local CLI backend. - Backend-specific smoke defaults are defined within the owning extension’s
cli-backend.tsfile. - Enable this by running pnpm test:live and setting
OPENCLAW_LIVE_CLI_BACKEND=1. - The default provider/model for this test is
claude-cli/claude-sonnet-4-6. - You can use optional overrides like
OPENCLAW_LIVE_CLI_BACKEND_MODELorOPENCLAW_LIVE_CLI_BACKEND_COMMANDto point to specific binaries. - Configure image probes with
OPENCLAW_LIVE_CLI_BACKEND_IMAGE_PROBE=1and specify the argument style withOPENCLAW_LIVE_CLI_BACKEND_IMAGE_ARG. - Control how image arguments are passed using
OPENCLAW_LIVE_CLI_BACKEND_IMAGE_MODE, such as “repeat” or “list”. - Validate resume flows with
OPENCLAW_LIVE_CLI_BACKEND_RESUME_PROBE=1to send a second turn. - Use
OPENCLAW_LIVE_CLI_BACKEND_MODEL_SWITCH_PROBE=0to disable the default Claude Sonnet to Opus continuity probe.
Example command:
OPENCLAW_LIVE_CLI_BACKEND=1 \ OPENCLAW_LIVE_CLI_BACKEND_MODEL="codex-cli/gpt-5.4" \ pnpm test:live src/gateway/gateway-cli-backend.live.test.tsDocker recipe:
pnpm test:docker:live-cli-backendSingle-provider Docker recipes:
pnpm test:docker:live-cli-backend:claudepnpm test:docker:live-cli-backend:claude-subscriptionpnpm test:docker:live-cli-backend:codexpnpm test:docker:live-cli-backend:geminiImportant notes for Docker users:
- The Docker runner is located at
scripts/test-live-cli-backend-docker.sh. - It runs the smoke test inside the repo Docker image as the non-root
nodeuser. - It installs matching Linux CLI packages like
@anthropic-ai/claude-codeor@openai/codexinto a cached directory. - The
claude-subscriptionlane requires portable Claude Code subscription OAuth through a credentials JSON orCLAUDE_CODE_OAUTH_TOKEN. - This subscription lane disables MCP and image probes by default because Claude routes third-party usage through extra-usage billing.
- The smoke test exercises the same end-to-end flow for Claude, Codex, and Gemini, including text turns, image classification, and MCP
crontool calls. - Claude’s default smoke also patches the session from Sonnet to Opus to verify the resumed session remembers previous notes.
Test OpenClaw ACP Bind Flows with Live Agents
Section titled “Test OpenClaw ACP Bind Flows with Live Agents”The OpenClaw ACP bind flow allows you to validate real conversation binding with live agents. This ensures that follow-up messages correctly land in the bound session transcript across different platforms.
- Run the test at
src/gateway/gateway-acp-bind.live.test.tsto validate the bind flow. - Enable the test by setting
OPENCLAW_LIVE_ACP_BIND=1. - The test sends a command like
/acp spawn <agent> --bind hereto bind a synthetic message-channel conversation. - It then sends a follow-up on that same conversation and verifies it lands in the bound ACP session transcript.
- Default agents in Docker include
claude,codex, andgemini, while the direct test defaults toclaude. - You can override the agent using
OPENCLAW_LIVE_ACP_BIND_AGENTor provide a custom command withOPENCLAW_LIVE_ACP_BIND_AGENT_COMMAND. - This lane uses the Gateway
chat.sendsurface with admin-only synthetic fields to attach message-channel context without external delivery. - When no custom command is set, the test uses the embedded
acpxplugin’s built-in agent registry.
Example command:
OPENCLAW_LIVE_ACP_BIND=1 \ OPENCLAW_LIVE_ACP_BIND_AGENT=claude \ pnpm test:live src/gateway/gateway-acp-bind.live.test.tsDocker recipe:
pnpm test:docker:live-acp-bindSingle-agent Docker recipes:
pnpm test:docker:live-acp-bind:claudepnpm test:docker:live-acp-bind:codexpnpm test:docker:live-acp-bind:geminiDocker implementation notes:
- The Docker runner is found at
scripts/test-live-acp-bind-docker.sh. - By default, it runs the ACP bind smoke against all supported live CLI agents in sequence.
- Use
OPENCLAW_LIVE_ACP_BIND_AGENTSto narrow the matrix to specific agents. - It stages CLI auth material into the container and installs
acpxinto a writable npm prefix. - Inside Docker, the runner sets
OPENCLAW_LIVE_ACP_BIND_ACPX_COMMANDso thatacpxkeeps provider environment variables available to the child CLI.
Run Codex app-server use smoke tests
Section titled “Run Codex app-server use smoke tests”You can use the OpenClaw Codex use to validate that your plugin-owned setup works correctly through the standard Gateway. This process ensures that the bundled codex plugin loads and that session threads can resume without issues.
- Goal: validate the plugin-owned Codex use through the normal Gateway using the
agentmethod. - Load the bundled
codexplugin. - Select
OPENCLAW_AGENT_RUNTIME=codex. - Send a first Gateway agent turn to
codex/gpt-5.4. - Send a second turn to the same OpenClaw session and verify the app-server thread can resume.
- Run
/codex statusand/codex modelsthrough the same Gateway command path. - Test file:
src/gateway/gateway-codex-use.live.test.ts. - Enable with:
OPENCLAW_LIVE_CODEX_HARNESS=1. - Default model:
codex/gpt-5.4. - Optional image probe:
OPENCLAW_LIVE_CODEX_HARNESS_IMAGE_PROBE=1. - Optional MCP/tool probe:
OPENCLAW_LIVE_CODEX_HARNESS_MCP_PROBE=1. - The smoke sets
OPENCLAW_AGENT_HARNESS_FALLBACK=noneso a broken Codex use cannot pass by silently falling back to PI. - Auth:
OPENAI_API_KEYfrom the shell/profile, plus optional copied~/.codex/auth.jsonand~/.codex/config.toml.
Local recipe:
source ~/.profileOPENCLAW_LIVE_CODEX_HARNESS=1 \ OPENCLAW_LIVE_CODEX_HARNESS_IMAGE_PROBE=1 \ OPENCLAW_LIVE_CODEX_HARNESS_MCP_PROBE=1 \ OPENCLAW_LIVE_CODEX_HARNESS_MODEL=codex/gpt-5.4 \ pnpm test:live -- src/gateway/gateway-codex-harness.live.test.tsDocker recipe:
source ~/.profilepnpm test:docker:live-codex-harnessDocker notes:
- The Docker runner lives at
scripts/test-live-codex-use-docker.sh. - It sources the mounted
~/.profile, passesOPENAI_API_KEY, copies Codex CLI auth files when present, installs@openai/codexinto a writable mounted npm prefix, stages the source tree, then runs only the Codex-use live test. - Docker enables the image and MCP/tool probes by default. Set `OPENCLAW_LIVE_CODEX_HARNESS_IMAGE
Run OpenClaw BytePlus coding plan live tests
Section titled “Run OpenClaw BytePlus coding plan live tests”Testing your OpenClaw BytePlus coding plan ensures that your integration works perfectly in a live environment. You can verify your setup by running the dedicated live test suite against the actual API.
- You can find the specific test file at
src/agents/byteplus.live.test.ts. - To enable and run the test, use the following command in your CLI:
BYTEPLUS_API_KEY=... BYTEPLUS_LIVE_TEST=1 pnpm test:live src/agents/byteplus.live.test.ts- If you need to override the default model, you can optionally set the following environment variable:
BYTEPLUS_CODING_MODEL=ark-code-latestTest ComfyUI workflow media live
Section titled “Test ComfyUI workflow media live”You can verify your ComfyUI workflows to confirm that image, video, and music generation paths are functioning as expected. This is especially helpful after you make changes to workflow submissions, polling mechanisms, or plugin registrations.
- The test file for this suite is located at
extensions/comfy/comfy.live.test.ts. - Execute the live test by running this command:
OPENCLAW_LIVE_TEST=1 COMFY_LIVE_TEST=1 pnpm test:live -- extensions/comfy/comfy.live.test.ts- This test covers several specific areas:
- It exercises the bundled Comfy image, video, and
music_generatepaths. - It will skip each capability automatically unless you have configured
models.providers.comfy.<capability>. - It is a great tool to use after you update Comfy workflow downloads or polling logic.
Execute Image generation live tests
Section titled “Execute Image generation live tests”OpenClaw provides a strong way to probe every registered image generation provider to ensure they are responsive and correctly configured. The test runner is designed to pick up your credentials directly from your environment to simplify the process.
- To run the main image generation test, use this command:
pnpm test:live src/image-generation/runtime.live.test.ts- You can also use the media-specific use:
pnpm test:live:media image- The scope of this test includes:
- It enumerates every single registered image-generation provider plugin in your setup.
- It automatically loads any missing provider environment variables from your login shell (such as
~/.profile) before starting the probe. - It prioritizes live environment API keys over stored auth profiles, ensuring that old keys in
auth-profiles.jsondon’t hide your actual shell credentials. - It skips any providers that do not have a usable auth configuration, profile, or model.
- It runs standard image-generation variants through the shared runtime, including
google:flash-generate,google:pro-generate,google:pro-edit, andopenai:default-generate.
- Currently, the bundled providers covered by this test are
openaiandgoogle. - If you want to narrow down the tests, you can use these optional environment variables:
OPENCLAW_LIVE_IMAGE_GENERATION_PROVIDERS="openai,google"OPENCLAW_LIVE_IMAGE_GENERATION_MODELS="openai/gpt-image-1,google/gemini-3.1-flash-image-preview"OPENCLAW_LIVE_IMAGE_GENERATION_CASES="google:flash-generate,google:pro-edit"- If you want to force the system to use profile-store keys and ignore environment overrides, set this flag:
OPENCLAW_LIVE_REQUIRE_PROFILE_KEYS=1Verify Music generation live functionality
Section titled “Verify Music generation live functionality”Validating your music generation setup ensures that providers like Google and MiniMax are ready to handle requests. This test suite checks the shared provider paths and supports both standard generation and editing features.
- Enable the music generation live tests with this command:
OPENCLAW_LIVE_TEST=1 pnpm test:live -- extensions/music-generation-providers.live.test.ts- Alternatively, you can use the dedicated media use:
pnpm test:live:media music- The test scope covers the following:
- It exercises the shared bundled music-generation provider path.
- It currently includes support for Google and MiniMax.
- It loads provider environment variables from your login shell (
~/.profile) before probing. - It uses live API keys from your environment ahead of stored auth profiles, so stale keys in
auth-profiles.jsonwon’t interfere. - It skips providers that lack a usable auth profile or model.
- It runs both runtime modes when available:
generatefor prompt-only inputs, andeditif the provider hascapabilities.edit.enabledset to true.
- The current shared-lane coverage includes
google(for bothgenerateandedit) andminimax(forgenerate). Note that ComfyUI uses a separate live file and is not part of this specific sweep. - You can narrow the music tests using these variables:
OPENCLAW_LIVE_MUSIC_GENERATION_PROVIDERS="google,minimax"OPENCLAW_LIVE_MUSIC_GENERATION_MODELS="google/lyria-3-clip-preview,minimax/music-2.5+"- To force the use of profile-store keys instead of environment overrides, use:
OPENCLAW_LIVE_REQUIRE_PROFILE_KEYS=1Run OpenClaw Video Generation Live Tests
Section titled “Run OpenClaw Video Generation Live Tests”Testing your OpenClaw video generation setup ensures that your API integrations are functioning correctly with live providers. You can use these tests to verify everything from basic text-to-video requests to advanced transformations across different platforms.
- You can find the main test file at
extensions/video-generation-providers.live.test.ts. - To enable the live test, run this command in your CLI:
OPENCLAW_LIVE_TEST=1 pnpm test:live -- extensions/video-generation-providers.live.test.ts- You can also use the dedicated testing tool for video specifically:
pnpm test:live:media video- This process exercises the shared bundled video-generation provider path and defaults to a release-safe smoke path. It uses non-FAL providers, one text-to-video request per provider, and a one-second lobster prompt.
- The system applies a per-provider operation cap defined by
OPENCLAW_LIVE_VIDEO_GENERATION_TIMEOUT_MS, which is180000by default. - FAL is skipped by default because provider-side queue latency can dominate release time. If you want to run it explicitly, pass
--video-providers falor setOPENCLAW_LIVE_VIDEO_GENERATION_PROVIDERS="fal". - The test loads provider environment variables from your login shell (
~/.profile) before probing the system. - It uses live/env API keys ahead of stored auth profiles by default. This ensures that stale test keys in
auth-profiles.jsondo not mask your real shell credentials. - Any providers that do not have a usable auth, profile, or model will be skipped automatically.
- By default, the test only runs the
generatemode. - If you want to run declared transform modes like
imageToVideoorvideoToVideo, setOPENCLAW_LIVE_VIDEO_GENERATION_FULL_MODES=1. - Note that
imageToVideois currently skipped in the shared sweep forvydrabecause the bundledveo3is text-only and the bundledklingrequires a remote image URL. - For specific Vydra coverage, you can run:
OPENCLAW_LIVE_TEST=1 OPENCLAW_LIVE_VYDRA_VIDEO=1 pnpm test:live -- extensions/vydra/vydra.live.test.tsThis file runs veo3 text-to-video and a kling lane using a remote image JSON fixture.
14. Currently, videoToVideo live coverage is limited to runway when the selected model is runway/gen4_aleph.
15. Several providers are skipped for videoToVideo in the shared sweep, including alibaba, qwen, and xai because they require remote http(s) or MP4 reference URLs.
16. google is skipped because the current shared Gemini/Veo lane uses local buffer-backed input which is not accepted in the shared sweep.
17. openai is also skipped because the current shared lane lacks organization-specific video inpaint or remix access guarantees.
18. You can narrow your testing scope using these optional environment variables:
OPENCLAW_LIVE_VIDEO_GENERATION_PROVIDERS="google,openai,runway"OPENCLAW_LIVE_VIDEO_GENERATION_MODELS="google/veo-3.1-fast-generate-preview,openai/sora-2,runway/gen4_aleph"OPENCLAW_LIVE_VIDEO_GENERATION_SKIP_PROVIDERS=""to include every provider, including FAL.OPENCLAW_LIVE_VIDEO_GENERATION_TIMEOUT_MS=60000to reduce the timeout for a faster smoke run.
- If you need to force profile-store auth and ignore environment overrides, set
OPENCLAW_LIVE_REQUIRE_PROFILE_KEYS=1.
Test Media Using the OpenClaw Live Utility
Section titled “Test Media Using the OpenClaw Live Utility”The OpenClaw media testing utility provides a unified way to run live suites for different media types through a single repo-native entrypoint. It simplifies your workflow by managing environment variables and provider selection automatically for image, music, video, and other assets.
- Use the following command to run the shared live suites:
pnpm test:live:media- This utility automatically loads any missing provider environment variables from your
~/.profile. - It narrows each suite down to providers that currently have usable auth by default, saving you time.
- Because it reuses
scripts/test-live.mjs, the heartbeat and quiet-mode behavior remain consistent with other tests. - You can run the full suite or target specific areas with these examples:
- To run everything:
pnpm test:live:media- To test specific providers for image and video:
pnpm test:live:media image video --providers openai,google,minimax- To test video for specific providers while including all others:
pnpm test:live:media video --video-providers openai,runway --all-providers- To run music tests in quiet mode:
pnpm test:live:media music --quietRun OpenClaw tests with Docker runners
Section titled “Run OpenClaw tests with Docker runners”When you want to ensure your code works perfectly in a Linux environment, OpenClaw Docker runners are the way to go. These tools allow you to run tests in isolated containers, keeping your local machine clean while verifying complex integrations.
- Live-model runners like
test:docker:live-modelsandtest:docker:live-gatewayrun specific profile-key live files inside the repo Docker image. They targetsrc/agents/models.profiles.live.test.tsandsrc/gateway/gateway-models.profiles.live.test.ts, mounting your local config directory and workspace. - These runners default to a smaller smoke cap to keep things fast. For example,
test:docker:live-modelsdefaults toOPENCLAW_LIVE_MAX_MODELS=12. - You can override environment variables like
OPENCLAW_LIVE_GATEWAY_SMOKE=1,OPENCLAW_LIVE_GATEWAY_MAX_MODELS=8,OPENCLAW_LIVE_GATEWAY_STEP_TIMEOUT_MS=45000, andOPENCLAW_LIVE_GATEWAY_MODEL_TIMEOUT_MS=90000if you need a more exhaustive scan. - The
test:docker:allcommand builds the live Docker image once usingtest:docker:live-buildand then reuses it for both live lanes. - Container smoke runners such as
test:docker:openwebui,test:docker:onboard,test:docker:gateway-network,test:docker:mcp-channels, andtest:docker:pluginsboot real containers to verify high-level integration paths.
The live-model Docker runners handle CLI authentication by bind-mounting the necessary homes and copying them into the container. This allows external CLI OAuth to refresh tokens without changing your host’s authentication store.
- Direct models:
pnpm test:docker:live-models(script: scripts/test-live-models-docker.sh)
- ACP bind smoke:
pnpm test:docker:live-acp-bind(script: scripts/test-live-acp-bind-docker.sh)
- CLI backend smoke:
pnpm test:docker:live-cli-backend(script: scripts/test-live-cli-backend-docker.sh)
- Codex app-server use smoke:
pnpm test:docker:live-codex-harness(script: scripts/test-live-codex-use-docker.sh)
- Gateway + dev agent:
pnpm test:docker:live-gateway(script: scripts/test-live-gateway-models-docker.sh)
- Open WebUI live smoke:
pnpm test:docker:openwebui(script: scripts/e2e/openwebui-docker.sh)
- Onboarding wizard (TTY, full scaffolding):
pnpm test:docker:onboard(script: scripts/e2e/onboard-docker.sh)
- Gateway networking (two containers, WS auth + health):
pnpm test:docker:gateway-network(script: scripts/e2e/gateway-network-docker.sh)
- MCP channel bridge:
pnpm test:docker:mcp-channels(script: scripts/e2e/mcp-channels-docker.sh)
- Plugins smoke:
pnpm test:docker:plugins(script: scripts/e2e/plugins-docker.sh)
To keep the runtime image slim, these runners bind-mount your current checkout as read-only and stage it into a temporary directory inside the container. This process skips large caches like .pnpm-store, .worktrees, __openclaw_vitest__, and build outputs so you don’t waste time copying unnecessary files.
- The runners set
OPENCLAW_SKIP_CHANNELS=1to prevent starting real Telegram or Discord workers inside the container. - If you use
test:docker:openwebui, it starts an OpenClaw Gateway container with OpenAI-compatible API endpoints and verifies the connection through a real chat request. - Successful runs for Open WebUI will print a JSON payload like
{ "ok": true, "model": "openclaw/default", ... }. - The
test:docker:mcp-channelsrunner is deterministic and doesn’t need real accounts; it verifies the stdio MCP bridge by inspecting raw frames directly.
If you need to perform a manual ACP plain-language thread smoke (not for CI), you can use this script:
bun scripts/dev/discord-acp-plain-language-smoke.ts --channel <discord-channel-id> ...You should keep this script for debugging workflows, especially for validating ACP thread routing. Here are some useful environment variables you can use:
OPENCLAW_CONFIG_DIR=...(default:~/.openclaw) mounted to/home/node/.openclaw.OPENCLAW_WORKSPACE_DIR=...(default:~/.openclaw/workspace) mounted to/home/node/.openclaw/workspace.OPENCLAW_PROFILE_FILE=...(default:~/.profile) mounted to/home/node/.profileand sourced before running tests.OPENCLAW_DOCKER_PROFILE_ENV_ONLY=1to verify only env vars sourced fromOPENCLAW_PROFILE_FILE, using temporary config/workspace dirs and no external CLI auth mounts.OPENCLAW_DOCKER_CLI_TOOLS_DIR=...(default:~/.cache/openclaw/docker-cli-tools) mounted to/home/node/.npm-globalfor cached CLI installs inside Docker.- External CLI auth dirs/files under
$HOMEare mounted read-only under/host-auth..., then copied into/home/node/...before tests start. This includes default dirs like.minimaxand files like~/.codex/auth.json,~/.codex/config.toml,.claude.json,~/.claude/.credentials.json,~/.claude/settings.json, and~/.claude/settings.local.json. - Narrowed provider runs mount only the needed dirs/files inferred from
OPENCLAW_LIVE_PROVIDERS/OPENCLAW_LIVE_GATEWAY_PROVIDERS. - Override manually with
OPENCLAW_DOCKER_AUTH_DIRS=all,OPENCLAW_DOCKER_AUTH_DIRS=none, or a comma list likeOPENCLAW_DOCKER_AUTH_DIRS=.claude,.codex. OPENCLAW_LIVE_GATEWAY_MODELS=.../OPENCLAW_LIVE_MODELS=...to narrow the run.OPENCLAW_LIVE_GATEWAY_PROVIDERS=.../OPENCLAW_LIVE_PROVIDERS=...to filter providers in-container.OPENCLAW_SKIP_DOCKER_BUILD=1to reuse an existingopenclaw:local-liveimage for reruns that do not need a rebuild.OPENCLAW_LIVE_REQUIRE_PROFILE_KEYS=1to ensure creds come from the profile store (not env).OPENCLAW_OPENWEBUI_MODEL=...to choose the model exposed by the Gateway for the Open WebUI smoke.OPENCLAW_OPENWEBUI_PROMPT=...to override the nonce-check prompt used by the Open WebUI smoke.OPENWEBUI_IMAGE=...to override the pinned Open WebUI image tag.
Validate documentation with Docs sanity checks
Section titled “Validate documentation with Docs sanity checks”Maintaining high-quality documentation is just as important as writing good code. You can use these built-in checks to ensure your edits don’t introduce broken links or formatting issues in the OpenClaw docs.
- Run basic documentation checks:
pnpm check:docs- Run full Mintlify anchor validation for in-page heading checks:
pnpm docs:check-links:anchorsRun Offline Regression Tests for OpenClaw
Section titled “Run Offline Regression Tests for OpenClaw”You can perform “real pipeline” regressions without needing actual providers by using these CI-safe tests. These tests help you verify that the OpenClaw Gateway and agent loops are functioning as expected in your local environment.
- Test the Gateway tool calling by using a mock OpenAI setup alongside the real Gateway and agent loop in
src/gateway/gateway.test.ts. You should look for the specific case: “runs a mock OpenAI tool call end-to-end via gateway agent loop”. - Validate the Gateway wizard by checking the WS
wizard.startandwizard.nextfunctions insrc/gateway/gateway.test.ts. This ensures the system correctly writes configuration files and enforces authentication, specifically in the case: “runs wizard over ws and writes auth token config”.
Improve Agent Reliability and Skills Evaluation
Section titled “Improve Agent Reliability and Skills Evaluation”You already have access to several CI-safe tests that function as reliability evaluations for your agent. These tests ensure that the core logic of your OpenClaw setup remains stable as you add new features.
- Use mock tool-calling through the real Gateway and agent loop as defined in
src/gateway/gateway.test.ts. - Run end-to-end wizard flows that validate how sessions are wired together and how configuration changes impact the system in
src/gateway/gateway.test.ts.
There are still a few things you need to handle for Skills to be fully tested:
- Decisioning: You need to confirm that when skills are listed in a prompt, the agent chooses the correct one and ignores skills that do not apply.
- Compliance: You must verify that the agent reads the
SKILL.mdfile before taking action and follows all the required steps and arguments. - Tool Order and History: Assert the correct order of tools in multi-turn scenarios and ensure session history carries over.
- Sandbox Boundaries: Confirm that the agent respects the defined boundaries of the sandbox.
Your future evaluations should focus on being deterministic first:
- Build a scenario runner using mock providers to assert tool calls,
SKILL.mdfile reads, and session wiring while testing usage, gating, and prompt injection. - Deploy optional live evaluations that are environment-gated only after your CI-safe suite is fully operational.
Master OpenClaw contract tests for plugins and channels
Section titled “Master OpenClaw contract tests for plugins and channels”OpenClaw contract tests ensure that every plugin and channel you register follows its interface contract perfectly. These tests iterate through all discovered plugins and run a suite of assertions to check both their shape and their behavior.
The default pnpm test unit lane skips these shared files on purpose. You should run these contract commands manually when you modify shared channel or provider surfaces:
- Run all contracts using pnpm test:contracts.
- Run channel contracts only with pnpm test:contracts:channels.
- Run provider contracts only using pnpm test:contracts:plugins.
You can find channel contracts in src/channels/plugins/contracts/*.contract.test.ts. They cover the following areas:
- plugin - Basic plugin shape including id, name, capabilities, and metadata.
- setup - The setup wizard contract.
- session-binding - How sessions are bound.
- outbound-payload - The structure of message payloads.
- inbound - Handling of inbound messages.
- actions - Handlers for channel actions.
- threading - How thread IDs are handled.
- directory - The directory and roster API.
- group-policy - How group policies are enforced.
Provider status contracts are located in src/plugins/contracts/*.contract.test.ts:
- status - Probes for channel status.
- registry - The shape of the plugin registry.
General provider contracts are also found in src/plugins/contracts/*.contract.test.ts:
- auth - The auth flow contract.
- auth-choice - Selection of auth choices.
- catalog - The model catalog API.
- discovery - How plugins are discovered.
- loader - How plugins are loaded.
- runtime - The provider runtime.
- shape - The interface and shape of the plugin.
- wizard - The setup wizard.
You should run these tests in these specific scenarios:
- After you change plugin-sdk exports or subpaths.
- After you add or change a channel or provider plugin.
- After you refactor how plugins are registered or discovered.
- When you update any shared plugin logic.
These tests run in CI and do not require real API keys to function.
Prevent bugs with OpenClaw regression testing
Section titled “Prevent bugs with OpenClaw regression testing”When you fix a provider or model issue discovered in a live environment, you should add a regression test to stop it from happening again. This keeps the OpenClaw codebase stable and reliable as you add more features.
- Add a CI-safe regression whenever you can. You can use a mock or stub provider, or just capture the exact request-shape transformation to verify the fix.
- If the bug only happens in live environments due to things like rate limits or auth policies, keep the test narrow. Make it opt-in using environment variables so it does not impact standard test runs.
- Try to hit the smallest layer possible to catch the bug. If it is a provider request conversion or replay bug, use a direct models test. If it is a Gateway session, history, tool pipeline, or state bug, use a Gateway live smoke test or a CI-safe Gateway mock test.
- Ensure the regression test is documented within the test suite.
You also need to be aware of the SecretRef traversal guardrail. The test in src/secrets/exec-secret-ref-id-parity.test.ts gets a sampled target for each SecretRef class from registry metadata using listSecretTargetRegistryEntries(). It then checks that traversal-segment exec ids are blocked.
- If you add a new
includeInPlanSecretRef target family insrc/secrets/target-registry-data.ts, you must updateclassifyTargetClassin that test. - The test will fail on purpose if it sees unclassified target ids so you do not skip new classes by accident.
OpenClaw Expert
Still stuck?
If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.