Skip to content

Master OpenClaw Testing: Run Unit, E2E, and Live Suites

OpenClaw testing is designed to be efficient and reliable, whether you are doing a quick check or a full validation of your changes. You can use these commands to maintain high code quality during your daily development workflow and ensure your Gateway is always performing as expected.

  1. Run the full gate, which is expected before you push any code: pnpm build && pnpm check && pnpm check:test-types && pnpm test
  2. Execute a faster local full-suite run if you are working on a machine with plenty of resources: pnpm test:max
  3. Use the direct Vitest watch loop for active development: pnpm test:watch
  4. Target specific files directly, which now also routes extension and channel paths correctly: pnpm test extensions/discord/src/monitor/message-handler.preflight.test.ts
  5. Iterate on a single failure by using targeted runs first before running the entire suite.
  6. Start the Docker-backed QA site for manual testing: pnpm qa:lab:up
  7. Run the Linux VM-backed QA lane using Multipass: pnpm openclaw qa suite --runner multipass --scenario channel-chat-baseline

When you touch tests or want extra confidence in your changes:

  1. Run the coverage gate to check your code coverage: pnpm test:coverage
  2. Execute the E2E suite for integration testing: pnpm test:e2e

When you are debugging real providers or models, which requires real credentials:

  1. Run the live suite for models and Gateway tool or image probes: pnpm test:live
  2. Target a single live file quietly to focus on one issue: pnpm test:live -- src/agents/models.profiles.live.test.ts

Tip: When you only need to fix one failing case, you should prefer narrowing down the live tests using the allowlist environment variables described below.

These commands provide a higher level of realism for your QA-lab work by simulating actual environments and transport layers. They sit alongside the main test suites to ensure your Gateway handles real-world scenarios correctly across different platforms.

  1. pnpm openclaw qa suite: This runs repo-backed QA scenarios directly on your host. It runs multiple selected scenarios in parallel by default with isolated Gateway workers. The qa-channel defaults to a concurrency of 4, which is bounded by the selected scenario count. You can tune the worker count with --concurrency <count> or use --concurrency 1 for the older serial lane. It exits with a non-zero code if any scenario fails, unless you use --allow-failures when you want artifacts without a failing exit code. It supports provider modes like live-frontier, mock-openai, and aimock. The aimock mode starts a local AIMock-backed provider server for experimental fixture and protocol-mock coverage.
  2. pnpm openclaw qa suite --runner multipass: This runs the same QA suite inside a disposable Multipass Linux VM. It keeps the same scenario-selection behavior as the host-based qa suite and reuses the same provider/model selection flags. Live runs forward the supported QA auth inputs that are practical for the guest, such as environment-based provider keys, the QA live provider config path, and CODEX_HOME when present. Output directories must stay under the repo root so the guest can write back through the mounted workspace. It writes the normal QA report, summary, and Multipass logs under .artifacts/qa-e2e/....
  3. pnpm qa:lab:up: This starts the Docker-backed QA site for operator-style QA work.
  4. pnpm openclaw qa aimock: This starts only the local AIMock provider server for direct protocol smoke testing.
  5. pnpm openclaw qa matrix: This runs the Matrix live QA lane against a disposable Docker-backed Tuwunel homeserver. This QA host is currently for repo and development use only, as packaged OpenClaw installs do not ship qa-lab. It provisions three temporary Matrix users (driver, sut, observer) and one private room, then starts a QA Gateway child with the real Matrix plugin as the SUT transport. It uses the pinned stable Tuwunel image ghcr.io/matrix-construct/tuwunel:v1.5.1 by default, which you can override with OPENCLAW_QA_MATRIX_TUWUNEL_IMAGE. It writes a Matrix QA report, summary, observed-events artifact, and combined logs under .artifacts/qa-e2e/....
  6. pnpm openclaw qa telegram: This runs the Telegram live QA lane against a real private group using the driver and SUT bot tokens from your environment. It requires OPENCLAW_QA_TELEGRAM_GROUP_ID, OPENCLAW_QA_TELEGRAM_DRIVER_BOT_TOKEN, and OPENCLAW_QA_TELEGRAM_SUT_BOT_TOKEN. It supports --credential-source convex for shared pooled credentials. It requires two distinct bots in the same private group, with the SUT bot exposing a Telegram username. For stable bot-to-bot observation, you should enable Bot-to-Bot Communication Mode in @BotFather for both bots. It writes a Telegram QA report, summary, and observed-messages artifact under .artifacts/qa-e2e/....

Live transport lanes share one standard contract so new transports do not drift:

qa-channel remains the broad synthetic QA suite and is not part of the live transport coverage matrix.

LaneCanaryMention gatingAllowlist blockTop-level replyRestart resumeThread follow-upThread isolationReaction observationHelp command
Matrixxxxxxxxx
Telegramxx

Shared Telegram credentials via Convex (v1)

Section titled “Shared Telegram credentials via Convex (v1)”

When you enable --credential-source convex (or set OPENCLAW_QA_CREDENTIAL_SOURCE=convex) for openclaw qa telegram, the QA lab acquires an exclusive lease from a Convex-backed pool. It heartbeats that lease while the lane is running and releases it upon shutdown.

Reference Convex project scaffold:

  1. qa/convex-credential-broker/

Required environment variables:

  1. OPENCLAW_QA_CONVEX_SITE_URL (for example https://your-deployment.convex.site)
  2. One secret for the selected role: OPENCLAW_QA_CONVEX_SECRET_MAINTAINER for maintainer or OPENCLAW_QA_CONVEX_SECRET_CI for ci.
  3. Credential role selection: Use the CLI flag --credential-role maintainer|ci or the environment default OPENCLAW_QA_CREDENTIAL_ROLE (defaults to ci in CI, maintainer otherwise).

Optional environment variables:

  1. OPENCLAW_QA_CREDENTIAL_LEASE_TTL_MS (default 1200000)
  2. OPENCLAW_QA_CREDENTIAL_HEARTBEAT_INTERVAL_MS (default 30000)
  3. OPENCLAW_QA_CREDENTIAL_ACQUIRE_TIMEOUT_MS (default 90000)
  4. OPENCLAW_QA_CREDENTIAL_HTTP_TIMEOUT_MS (default 15000)
  5. OPENCLAW_QA_CONVEX_ENDPOINT_PREFIX (default /qa-credentials/v1)
  6. OPENCLAW_QA_CREDENTIAL_OWNER_ID (optional trace id)
  7. OPENCLAW_QA_ALLOW_INSECURE_HTTP=1 allows loopback http:// Convex URLs for local-only development.

Maintainer admin commands for pool management require OPENCLAW_QA_CONVEX_SECRET_MAINTAINER specifically. You can use these CLI helpers:

Terminal window
pnpm openclaw qa credentials add --kind telegram --payload-file qa/telegram-credential.json
pnpm openclaw qa credentials list --kind telegram
pnpm openclaw qa credentials remove --credential-id <credential-id>

Use --json for machine-readable output in scripts and CI utilities. The default endpoint contract (OPENCLAW_QA_CONVEX_SITE_URL + /qa-credentials/v1) includes:

  1. POST /acquire: Request includes { kind, ownerId, actorRole, leaseTtlMs, heartbeatIntervalMs }. Success returns { status: "ok", credentialId, leaseToken, payload, leaseTtlMs?, heartbeatIntervalMs? }.
  2. POST /heartbeat: Request includes { kind, ownerId, actorRole, credentialId, leaseToken, leaseTtlMs }. Success returns { status: "ok" }.
  3. POST /release: Request includes { kind, ownerId, actorRole, credentialId, leaseToken }. Success returns { status: "ok" }.
  4. POST /admin/add: Request includes { kind, actorId, payload, note?, status? } (maintainer secret only).
  5. POST /admin/remove: Request includes { credentialId, actorId } (maintainer secret only).
  6. POST /admin/list: Request includes { kind?, status?, includePayload?, limit? } (maintainer secret only).

The JSON payload shape for the Telegram kind consists of:

  1. { groupId: string, driverToken: string, sutToken: string }
  2. groupId must be a numeric Telegram chat id string.
  3. admin/add validates this shape for kind: "telegram" and rejects malformed payloads.

Adding a channel to the markdown QA system requires exactly two things: a transport adapter for the channel and a scenario pack that exercises the channel contract. You should not add a new top-level QA command root when the shared qa-lab host can own the flow.

The qa-lab owns the shared host mechanics:

  1. The openclaw qa command root.
  2. Suite startup and teardown.
  3. Worker concurrency and artifact writing.
  4. Report generation and scenario execution.
  5. Compatibility aliases for older qa-channel scenarios.

Runner plugins own the transport contract:

  1. How openclaw qa <runner> is mounted beneath the shared qa root.
  2. How the Gateway is configured for that transport.
  3. How readiness is checked and how inbound events are injected.
  4. How outbound messages are observed and how transcripts are exposed.
  5. How transport-backed actions are executed and how cleanup is handled.

The minimum adoption bar for a new channel includes:

  1. Keep qa-lab as the owner of the shared qa root.
  2. Implement the transport runner on the shared qa-lab host seam.
  3. Keep transport-specific mechanics inside the runner plugin or channel use.
  4. Mount the runner as openclaw qa <runner> instead of registering a competing root command. Runner plugins should declare qaRunners in openclaw.plugin.json and export a matching qaRunnerCliRegistrations array from runtime-api.ts.
  5. Author or adapt markdown scenarios under the themed qa/scenarios/ directories.
  6. Use the generic scenario helpers for new scenarios.
  7. Keep existing compatibility aliases working unless the repo is doing an intentional migration.

Preferred generic helper names for new scenarios are:

  1. waitForTransportReady
  2. waitForChannelReady
  3. injectInboundMessage
  4. injectOutboundMessage
  5. waitForTransportOutboundMessage
  6. waitForChannelOutboundMessage
  7. waitForNoTransportOutbound
  8. getTransportSnapshot
  9. readTransportMessage
  10. readTransportTranscript
  11. formatTransportTranscript
  12. resetTransport

Compatibility aliases remain available for existing scenarios, including:

  1. waitForQaChannelReady
  2. waitForOutboundMessage
  3. waitForNoOutbound
  4. formatConversationTranscript
  5. resetBus

OpenClaw uses a tiered testing strategy where each suite offers increasing realism at the cost of speed and complexity. You should think of these suites as a progression that helps you catch everything from simple logic errors to complex provider integration issues.

  1. Command: pnpm test
  2. Config: Ten sequential shard runs (vitest.full-*.config.ts) over the existing scoped Vitest projects.
  3. Files: Core/unit inventories under src/**/*.test.ts, packages/**/*.test.ts, test/**/*.test.ts, and the whitelisted ui node tests.
  4. Scope: Pure unit tests, in-process integration tests (for Gateway auth, routing, tooling, parsing, and config), and deterministic regressions for known bugs.
  5. Expectations: These run in CI, require no real keys, and should be fast and stable.

Technical notes for the unit suite:

  1. Sharding: Untargeted pnpm test now runs eleven smaller shard configs (like core-unit-src, agentic, extensions) instead of one giant process. This cuts peak RSS on loaded machines and avoids resource starvation.
  2. Watch Mode: pnpm test --watch still uses the native root vitest.config.ts project graph because a multi-shard watch loop is not practical.
  3. Targeted Runs: pnpm test, pnpm test:watch, and pnpm test:perf:imports route explicit file targets through scoped lanes first to avoid the full root project startup tax.
  4. Changed Files: pnpm test:changed expands changed git paths into the same scoped lanes when the diff only touches routable source/test files. pnpm check:changed is the normal smart local gate that classifies the diff into core, extensions, apps, and more.
  5. Fast Lanes: Import-light unit tests from agents, commands, and plugin-sdk route through the unit-fast lane, which skips test/setup-openclaw-runtime.ts.
  6. Auto-reply: This now has three dedicated buckets to keep the heaviest reply use work off the cheaper status and token tests.
  7. Embedded Runner: When you change message-tool discovery inputs, you must keep both levels of coverage. Focused helper regressions and embedded runner integration suites (like src/agents/pi-embedded-runner/compact.hooks.test.ts) verify that behavior flows through the real paths.
  8. Pool and Isolation: Base Vitest config now defaults to threads and uses isolate: false across the root projects, e2e, and live configs. The shared scripts/run-vitest.mjs launcher adds --no-maglev for Vitest child Node.js processes by default to reduce V8 compile churn.
  9. Local Iteration: pnpm changed:lanes shows which architectural lanes a diff triggers. Local worker auto-scaling is conservative and backs off when the host load average is high.
  10. Performance Debugging: pnpm test:perf:imports enables Vitest import-duration reporting. pnpm test:perf:changed:bench -- --ref <git-ref> compares routed test:changed against the native root-project path. pnpm test:perf:profile:main writes a main-thread CPU profile for Vitest and Vite startup overhead.
  1. Command: pnpm test:e2e
  2. Config: vitest.e2e.config.ts
  3. Files: src/**/*.e2e.test.ts, test/**/*.e2e.test.ts
  4. Scope: Multi-instance Gateway end-to-end behavior, WebSocket and HTTP surfaces, node pairing, and heavier networking.
  5. Expectations: Runs in CI (when enabled), requires no real keys, but has more moving parts than unit tests.
  6. Overrides: Use OPENCLAW_E2E_WORKERS=<n> to force worker count (capped at 16) or OPENCLAW_E2E_VERBOSE=1 to re-enable verbose console output.
  1. Command: pnpm test:e2e:openshell
  2. File: test/openshell-sandbox.e2e.test.ts
  3. Scope: Starts an isolated OpenShell Gateway on the host via Docker, creates a sandbox from a temporary local Dockerfile, and exercises the OpenShell backend over real SSH.
  4. Expectations: This is opt-in only and requires a local openshell CLI plus a working Docker daemon.
  5. Overrides: Set OPENCLAW_E2E_OPENSHELL=1 to enable the test or OPENCLAW_E2E_OPENSHELL_COMMAND=/path/to/openshell to point at a specific binary.
  1. Command: pnpm test:live
  2. Config: vitest.live.config.ts
  3. Files: src/**/*.live.test.ts
  4. Scope: Verifies if a provider or model actually works today with real credentials. It catches provider format changes, tool-calling quirks, auth issues, and rate limit behavior.
  5. Expectations: Not CI-stable by design due to real network and provider policies. It costs money and uses rate limits, so you should prefer running narrowed subsets.
  6. Key Management: Live runs source ~/.profile to pick up missing API keys. By default, they isolate HOME and copy config into a temp test home. Set OPENCLAW_LIVE_USE_REAL_HOME=1 only if you intentionally need to use your real home directory.
  7. Quiet Mode: pnpm test:live defaults to a quieter mode, suppressing extra notices and Gateway bootstrap logs. Set OPENCLAW_LIVE_TEST_QUIET=0 for full logs.
  8. Key Rotation: Set *_API_KEYS with comma/semicolon format or *_API_KEY_1, *_API_KEY_2 for provider-specific rotation. Tests will retry on rate limit responses.
  9. Progress Output: Live suites emit progress lines to stderr so long provider calls are visibly active. You can tune direct-model heartbeats with OPENCLAW_LIVE_HEARTBEAT_MS and Gateway probes with OPENCLAW_LIVE_GATEWAY_HEARTBEAT_MS.

Deciding which test to run depends on the scope of your changes and the level of validation you need. Use this simple decision guide to pick the most efficient path for your current task.

  1. If you are editing logic or standard tests: Run pnpm test. If you have changed a significant amount of code, you should also run pnpm test:coverage.
  2. If you are touching Gateway networking, the WebSocket protocol, or node pairing: You should add pnpm test:e2e to your validation process.
  3. If you are debugging “my bot is down” issues, provider-specific failures, or complex tool calling: Run a narrowed subset of pnpm test:live to test against real models.

You can use the OpenClaw Android live test suite to verify that your connected mobile nodes are handling commands correctly. This ensures every advertised capability actually works as expected in a real-world environment.

  1. Locate the test file at src/gateway/android-node.capabilities.live.test.ts.
  2. Run the integration script using the command pnpm android:test:integration.
  3. The primary goal is to invoke every command currently advertised by a connected Android node and assert command contract behavior.
  4. Note that the scope involves preconditioned or manual setup, as the suite does not install, run, or pair the app automatically.
  5. The suite performs command-by-command Gateway node.invoke validation for your selected Android node.

To ensure the tests run correctly, you need to complete the following pre-setup:

  1. Confirm the Android app is already connected and paired to the Gateway.
  2. Keep the app in the foreground during the testing process.
  3. Grant all necessary permissions and capture consent for the capabilities you expect to pass.

You can also use optional target overrides to customize the test:

  1. Set OPENCLAW_ANDROID_NODE_ID or OPENCLAW_ANDROID_NODE_NAME to target a specific device.
  2. Use OPENCLAW_ANDROID_GATEWAY_URL, OPENCLAW_ANDROID_GATEWAY_TOKEN, or OPENCLAW_ANDROID_GATEWAY_PASSWORD to define connection parameters.
  3. Refer to the Android App documentation for full setup details.

Run OpenClaw Model Smoke Tests with Profile Keys

Section titled “Run OpenClaw Model Smoke Tests with Profile Keys”

Testing models in OpenClaw is split into two distinct layers to help you quickly identify if a failure is due to the provider’s API or the internal pipeline. This isolation makes debugging much easier when you are dealing with multiple LLM providers.

Layer 1: Direct model completion (no gateway)

Section titled “Layer 1: Direct model completion (no gateway)”
  1. Use the test at src/agents/models.profiles.live.test.ts to validate direct model access.
  2. The goal is to enumerate discovered models and use getApiKeyForModel to select models you have credentials for.
  3. This layer runs a small completion per model and includes targeted regressions where needed.
  4. Enable this by running pnpm test:live or setting OPENCLAW_LIVE_TEST=1 if you are invoking Vitest directly.
  5. You must set OPENCLAW_LIVE_MODELS=modern (or all) to actually run this suite; otherwise, it skips to keep pnpm test:live focused on Gateway smoke.
  6. Select models using OPENCLAW_LIVE_MODELS=modern for the modern allowlist, which includes Opus/Sonnet 4.6+, GPT-5.x + Codex, Gemini 3, GLM 4.7, MiniMax M2.7, and Grok 4.
  7. Use a comma-separated allowlist like OPENCLAW_LIVE_MODELS="openai/gpt-5.4,anthropic/claude-opus-4-6,..." for specific targets.
  8. Modern sweeps default to a curated high-signal cap, but you can set OPENCLAW_LIVE_MAX_MODELS=0 for an exhaustive sweep or a positive number for a smaller cap.
  9. Select specific providers using OPENCLAW_LIVE_PROVIDERS="google,google-antigravity,google-gemini-cli".
  10. By default, keys come from the profile store and env fallbacks, but you can set OPENCLAW_LIVE_REQUIRE_PROFILE_KEYS=1 to enforce profile store only.
  11. This layer exists to separate “provider API is broken or key is invalid” from “Gateway agent pipeline is broken” and contains small regressions like OpenAI or Codex reasoning replays.
  1. This test is located at src/gateway/gateway-models.profiles.live.test.ts and covers what “@openclaw” actually does.
  2. The goal is to spin up an in-process Gateway and create or patch an agent:dev:* session with a model override per run.
  3. It iterates through models-with-keys to assert a meaningful response without tools, a real tool invocation, and optional extra tool probes.
  4. The read probe involves the test writing a nonce file in the workspace and asking the agent to read it and echo the nonce back.
  5. The exec+read probe asks the agent to use exec to write a nonce into a temp file and then read it back.
  6. The image probe attaches a generated PNG containing a cat and randomized code, expecting the model to return “cat” plus the code.
  7. Implementation details are found in src/gateway/gateway-models.profiles.live.test.ts and src/gateway/live-image-probe.ts.
  8. Enable this layer using pnpm test:live.
  9. Select models via OPENCLAW_LIVE_GATEWAY_MODELS=all or a specific comma-separated list of “provider/model” strings.
  10. Use OPENCLAW_LIVE_GATEWAY_PROVIDERS to narrow the providers and avoid testing everything through OpenRouter.
  11. Tool and image probes are always active in this test; the image probe runs when the model advertises image input support.
  12. The high-level flow involves generating a tiny PNG, sending it via attachments, parsing it into images[] in the Gateway, and forwarding a multimodal message to the model.
  13. The assertion allows for minor OCR mistakes when the model replies with the code.

If you want to see which models you can test on your machine and their exact IDs, run these commands:

Terminal window
openclaw models list
openclaw models list --json

Validate OpenClaw CLI Backends for Claude and Codex

Section titled “Validate OpenClaw CLI Backends for Claude and Codex”

You can validate the OpenClaw Gateway and agent pipeline using local CLI backends without touching your default configuration. This is perfect for testing tools like Claude, Codex, or Gemini in a controlled environment.

  1. Use the test at src/gateway/gateway-cli-backend.live.test.ts to validate the pipeline with a local CLI backend.
  2. Backend-specific smoke defaults are defined within the owning extension’s cli-backend.ts file.
  3. Enable this by running pnpm test:live and setting OPENCLAW_LIVE_CLI_BACKEND=1.
  4. The default provider/model for this test is claude-cli/claude-sonnet-4-6.
  5. You can use optional overrides like OPENCLAW_LIVE_CLI_BACKEND_MODEL or OPENCLAW_LIVE_CLI_BACKEND_COMMAND to point to specific binaries.
  6. Configure image probes with OPENCLAW_LIVE_CLI_BACKEND_IMAGE_PROBE=1 and specify the argument style with OPENCLAW_LIVE_CLI_BACKEND_IMAGE_ARG.
  7. Control how image arguments are passed using OPENCLAW_LIVE_CLI_BACKEND_IMAGE_MODE, such as “repeat” or “list”.
  8. Validate resume flows with OPENCLAW_LIVE_CLI_BACKEND_RESUME_PROBE=1 to send a second turn.
  9. Use OPENCLAW_LIVE_CLI_BACKEND_MODEL_SWITCH_PROBE=0 to disable the default Claude Sonnet to Opus continuity probe.

Example command:

Terminal window
OPENCLAW_LIVE_CLI_BACKEND=1 \
OPENCLAW_LIVE_CLI_BACKEND_MODEL="codex-cli/gpt-5.4" \
pnpm test:live src/gateway/gateway-cli-backend.live.test.ts

Docker recipe:

Terminal window
pnpm test:docker:live-cli-backend

Single-provider Docker recipes:

Terminal window
pnpm test:docker:live-cli-backend:claude
pnpm test:docker:live-cli-backend:claude-subscription
pnpm test:docker:live-cli-backend:codex
pnpm test:docker:live-cli-backend:gemini

Important notes for Docker users:

  1. The Docker runner is located at scripts/test-live-cli-backend-docker.sh.
  2. It runs the smoke test inside the repo Docker image as the non-root node user.
  3. It installs matching Linux CLI packages like @anthropic-ai/claude-code or @openai/codex into a cached directory.
  4. The claude-subscription lane requires portable Claude Code subscription OAuth through a credentials JSON or CLAUDE_CODE_OAUTH_TOKEN.
  5. This subscription lane disables MCP and image probes by default because Claude routes third-party usage through extra-usage billing.
  6. The smoke test exercises the same end-to-end flow for Claude, Codex, and Gemini, including text turns, image classification, and MCP cron tool calls.
  7. Claude’s default smoke also patches the session from Sonnet to Opus to verify the resumed session remembers previous notes.

Test OpenClaw ACP Bind Flows with Live Agents

Section titled “Test OpenClaw ACP Bind Flows with Live Agents”

The OpenClaw ACP bind flow allows you to validate real conversation binding with live agents. This ensures that follow-up messages correctly land in the bound session transcript across different platforms.

  1. Run the test at src/gateway/gateway-acp-bind.live.test.ts to validate the bind flow.
  2. Enable the test by setting OPENCLAW_LIVE_ACP_BIND=1.
  3. The test sends a command like /acp spawn <agent> --bind here to bind a synthetic message-channel conversation.
  4. It then sends a follow-up on that same conversation and verifies it lands in the bound ACP session transcript.
  5. Default agents in Docker include claude, codex, and gemini, while the direct test defaults to claude.
  6. You can override the agent using OPENCLAW_LIVE_ACP_BIND_AGENT or provide a custom command with OPENCLAW_LIVE_ACP_BIND_AGENT_COMMAND.
  7. This lane uses the Gateway chat.send surface with admin-only synthetic fields to attach message-channel context without external delivery.
  8. When no custom command is set, the test uses the embedded acpx plugin’s built-in agent registry.

Example command:

Terminal window
OPENCLAW_LIVE_ACP_BIND=1 \
OPENCLAW_LIVE_ACP_BIND_AGENT=claude \
pnpm test:live src/gateway/gateway-acp-bind.live.test.ts

Docker recipe:

Terminal window
pnpm test:docker:live-acp-bind

Single-agent Docker recipes:

Terminal window
pnpm test:docker:live-acp-bind:claude
pnpm test:docker:live-acp-bind:codex
pnpm test:docker:live-acp-bind:gemini

Docker implementation notes:

  1. The Docker runner is found at scripts/test-live-acp-bind-docker.sh.
  2. By default, it runs the ACP bind smoke against all supported live CLI agents in sequence.
  3. Use OPENCLAW_LIVE_ACP_BIND_AGENTS to narrow the matrix to specific agents.
  4. It stages CLI auth material into the container and installs acpx into a writable npm prefix.
  5. Inside Docker, the runner sets OPENCLAW_LIVE_ACP_BIND_ACPX_COMMAND so that acpx keeps provider environment variables available to the child CLI.

You can use the OpenClaw Codex use to validate that your plugin-owned setup works correctly through the standard Gateway. This process ensures that the bundled codex plugin loads and that session threads can resume without issues.

  1. Goal: validate the plugin-owned Codex use through the normal Gateway using the agent method.
  2. Load the bundled codex plugin.
  3. Select OPENCLAW_AGENT_RUNTIME=codex.
  4. Send a first Gateway agent turn to codex/gpt-5.4.
  5. Send a second turn to the same OpenClaw session and verify the app-server thread can resume.
  6. Run /codex status and /codex models through the same Gateway command path.
  7. Test file: src/gateway/gateway-codex-use.live.test.ts.
  8. Enable with: OPENCLAW_LIVE_CODEX_HARNESS=1.
  9. Default model: codex/gpt-5.4.
  10. Optional image probe: OPENCLAW_LIVE_CODEX_HARNESS_IMAGE_PROBE=1.
  11. Optional MCP/tool probe: OPENCLAW_LIVE_CODEX_HARNESS_MCP_PROBE=1.
  12. The smoke sets OPENCLAW_AGENT_HARNESS_FALLBACK=none so a broken Codex use cannot pass by silently falling back to PI.
  13. Auth: OPENAI_API_KEY from the shell/profile, plus optional copied ~/.codex/auth.json and ~/.codex/config.toml.

Local recipe:

Terminal window
source ~/.profile
OPENCLAW_LIVE_CODEX_HARNESS=1 \
OPENCLAW_LIVE_CODEX_HARNESS_IMAGE_PROBE=1 \
OPENCLAW_LIVE_CODEX_HARNESS_MCP_PROBE=1 \
OPENCLAW_LIVE_CODEX_HARNESS_MODEL=codex/gpt-5.4 \
pnpm test:live -- src/gateway/gateway-codex-harness.live.test.ts

Docker recipe:

Terminal window
source ~/.profile
pnpm test:docker:live-codex-harness

Docker notes:

  1. The Docker runner lives at scripts/test-live-codex-use-docker.sh.
  2. It sources the mounted ~/.profile, passes OPENAI_API_KEY, copies Codex CLI auth files when present, installs @openai/codex into a writable mounted npm prefix, stages the source tree, then runs only the Codex-use live test.
  3. Docker enables the image and MCP/tool probes by default. Set `OPENCLAW_LIVE_CODEX_HARNESS_IMAGE

Run OpenClaw BytePlus coding plan live tests

Section titled “Run OpenClaw BytePlus coding plan live tests”

Testing your OpenClaw BytePlus coding plan ensures that your integration works perfectly in a live environment. You can verify your setup by running the dedicated live test suite against the actual API.

  1. You can find the specific test file at src/agents/byteplus.live.test.ts.
  2. To enable and run the test, use the following command in your CLI:
Terminal window
BYTEPLUS_API_KEY=... BYTEPLUS_LIVE_TEST=1 pnpm test:live src/agents/byteplus.live.test.ts
  1. If you need to override the default model, you can optionally set the following environment variable:
Terminal window
BYTEPLUS_CODING_MODEL=ark-code-latest

You can verify your ComfyUI workflows to confirm that image, video, and music generation paths are functioning as expected. This is especially helpful after you make changes to workflow submissions, polling mechanisms, or plugin registrations.

  1. The test file for this suite is located at extensions/comfy/comfy.live.test.ts.
  2. Execute the live test by running this command:
Terminal window
OPENCLAW_LIVE_TEST=1 COMFY_LIVE_TEST=1 pnpm test:live -- extensions/comfy/comfy.live.test.ts
  1. This test covers several specific areas:
  • It exercises the bundled Comfy image, video, and music_generate paths.
  • It will skip each capability automatically unless you have configured models.providers.comfy.<capability>.
  • It is a great tool to use after you update Comfy workflow downloads or polling logic.

OpenClaw provides a strong way to probe every registered image generation provider to ensure they are responsive and correctly configured. The test runner is designed to pick up your credentials directly from your environment to simplify the process.

  1. To run the main image generation test, use this command:
Terminal window
pnpm test:live src/image-generation/runtime.live.test.ts
  1. You can also use the media-specific use:
Terminal window
pnpm test:live:media image
  1. The scope of this test includes:
  • It enumerates every single registered image-generation provider plugin in your setup.
  • It automatically loads any missing provider environment variables from your login shell (such as ~/.profile) before starting the probe.
  • It prioritizes live environment API keys over stored auth profiles, ensuring that old keys in auth-profiles.json don’t hide your actual shell credentials.
  • It skips any providers that do not have a usable auth configuration, profile, or model.
  • It runs standard image-generation variants through the shared runtime, including google:flash-generate, google:pro-generate, google:pro-edit, and openai:default-generate.
  1. Currently, the bundled providers covered by this test are openai and google.
  2. If you want to narrow down the tests, you can use these optional environment variables:
Terminal window
OPENCLAW_LIVE_IMAGE_GENERATION_PROVIDERS="openai,google"
OPENCLAW_LIVE_IMAGE_GENERATION_MODELS="openai/gpt-image-1,google/gemini-3.1-flash-image-preview"
OPENCLAW_LIVE_IMAGE_GENERATION_CASES="google:flash-generate,google:pro-edit"
  1. If you want to force the system to use profile-store keys and ignore environment overrides, set this flag:
Terminal window
OPENCLAW_LIVE_REQUIRE_PROFILE_KEYS=1

Verify Music generation live functionality

Section titled “Verify Music generation live functionality”

Validating your music generation setup ensures that providers like Google and MiniMax are ready to handle requests. This test suite checks the shared provider paths and supports both standard generation and editing features.

  1. Enable the music generation live tests with this command:
Terminal window
OPENCLAW_LIVE_TEST=1 pnpm test:live -- extensions/music-generation-providers.live.test.ts
  1. Alternatively, you can use the dedicated media use:
Terminal window
pnpm test:live:media music
  1. The test scope covers the following:
  • It exercises the shared bundled music-generation provider path.
  • It currently includes support for Google and MiniMax.
  • It loads provider environment variables from your login shell (~/.profile) before probing.
  • It uses live API keys from your environment ahead of stored auth profiles, so stale keys in auth-profiles.json won’t interfere.
  • It skips providers that lack a usable auth profile or model.
  • It runs both runtime modes when available: generate for prompt-only inputs, and edit if the provider has capabilities.edit.enabled set to true.
  1. The current shared-lane coverage includes google (for both generate and edit) and minimax (for generate). Note that ComfyUI uses a separate live file and is not part of this specific sweep.
  2. You can narrow the music tests using these variables:
Terminal window
OPENCLAW_LIVE_MUSIC_GENERATION_PROVIDERS="google,minimax"
OPENCLAW_LIVE_MUSIC_GENERATION_MODELS="google/lyria-3-clip-preview,minimax/music-2.5+"
  1. To force the use of profile-store keys instead of environment overrides, use:
Terminal window
OPENCLAW_LIVE_REQUIRE_PROFILE_KEYS=1

Testing your OpenClaw video generation setup ensures that your API integrations are functioning correctly with live providers. You can use these tests to verify everything from basic text-to-video requests to advanced transformations across different platforms.

  1. You can find the main test file at extensions/video-generation-providers.live.test.ts.
  2. To enable the live test, run this command in your CLI:
Terminal window
OPENCLAW_LIVE_TEST=1 pnpm test:live -- extensions/video-generation-providers.live.test.ts
  1. You can also use the dedicated testing tool for video specifically:
Terminal window
pnpm test:live:media video
  1. This process exercises the shared bundled video-generation provider path and defaults to a release-safe smoke path. It uses non-FAL providers, one text-to-video request per provider, and a one-second lobster prompt.
  2. The system applies a per-provider operation cap defined by OPENCLAW_LIVE_VIDEO_GENERATION_TIMEOUT_MS, which is 180000 by default.
  3. FAL is skipped by default because provider-side queue latency can dominate release time. If you want to run it explicitly, pass --video-providers fal or set OPENCLAW_LIVE_VIDEO_GENERATION_PROVIDERS="fal".
  4. The test loads provider environment variables from your login shell (~/.profile) before probing the system.
  5. It uses live/env API keys ahead of stored auth profiles by default. This ensures that stale test keys in auth-profiles.json do not mask your real shell credentials.
  6. Any providers that do not have a usable auth, profile, or model will be skipped automatically.
  7. By default, the test only runs the generate mode.
  8. If you want to run declared transform modes like imageToVideo or videoToVideo, set OPENCLAW_LIVE_VIDEO_GENERATION_FULL_MODES=1.
  9. Note that imageToVideo is currently skipped in the shared sweep for vydra because the bundled veo3 is text-only and the bundled kling requires a remote image URL.
  10. For specific Vydra coverage, you can run:
Terminal window
OPENCLAW_LIVE_TEST=1 OPENCLAW_LIVE_VYDRA_VIDEO=1 pnpm test:live -- extensions/vydra/vydra.live.test.ts

This file runs veo3 text-to-video and a kling lane using a remote image JSON fixture. 14. Currently, videoToVideo live coverage is limited to runway when the selected model is runway/gen4_aleph. 15. Several providers are skipped for videoToVideo in the shared sweep, including alibaba, qwen, and xai because they require remote http(s) or MP4 reference URLs. 16. google is skipped because the current shared Gemini/Veo lane uses local buffer-backed input which is not accepted in the shared sweep. 17. openai is also skipped because the current shared lane lacks organization-specific video inpaint or remix access guarantees. 18. You can narrow your testing scope using these optional environment variables:

  • OPENCLAW_LIVE_VIDEO_GENERATION_PROVIDERS="google,openai,runway"
  • OPENCLAW_LIVE_VIDEO_GENERATION_MODELS="google/veo-3.1-fast-generate-preview,openai/sora-2,runway/gen4_aleph"
  • OPENCLAW_LIVE_VIDEO_GENERATION_SKIP_PROVIDERS="" to include every provider, including FAL.
  • OPENCLAW_LIVE_VIDEO_GENERATION_TIMEOUT_MS=60000 to reduce the timeout for a faster smoke run.
  1. If you need to force profile-store auth and ignore environment overrides, set OPENCLAW_LIVE_REQUIRE_PROFILE_KEYS=1.

Test Media Using the OpenClaw Live Utility

Section titled “Test Media Using the OpenClaw Live Utility”

The OpenClaw media testing utility provides a unified way to run live suites for different media types through a single repo-native entrypoint. It simplifies your workflow by managing environment variables and provider selection automatically for image, music, video, and other assets.

  1. Use the following command to run the shared live suites:
Terminal window
pnpm test:live:media
  1. This utility automatically loads any missing provider environment variables from your ~/.profile.
  2. It narrows each suite down to providers that currently have usable auth by default, saving you time.
  3. Because it reuses scripts/test-live.mjs, the heartbeat and quiet-mode behavior remain consistent with other tests.
  4. You can run the full suite or target specific areas with these examples:
  • To run everything:
Terminal window
pnpm test:live:media
  • To test specific providers for image and video:
Terminal window
pnpm test:live:media image video --providers openai,google,minimax
  • To test video for specific providers while including all others:
Terminal window
pnpm test:live:media video --video-providers openai,runway --all-providers
  • To run music tests in quiet mode:
Terminal window
pnpm test:live:media music --quiet

When you want to ensure your code works perfectly in a Linux environment, OpenClaw Docker runners are the way to go. These tools allow you to run tests in isolated containers, keeping your local machine clean while verifying complex integrations.

  1. Live-model runners like test:docker:live-models and test:docker:live-gateway run specific profile-key live files inside the repo Docker image. They target src/agents/models.profiles.live.test.ts and src/gateway/gateway-models.profiles.live.test.ts, mounting your local config directory and workspace.
  2. These runners default to a smaller smoke cap to keep things fast. For example, test:docker:live-models defaults to OPENCLAW_LIVE_MAX_MODELS=12.
  3. You can override environment variables like OPENCLAW_LIVE_GATEWAY_SMOKE=1, OPENCLAW_LIVE_GATEWAY_MAX_MODELS=8, OPENCLAW_LIVE_GATEWAY_STEP_TIMEOUT_MS=45000, and OPENCLAW_LIVE_GATEWAY_MODEL_TIMEOUT_MS=90000 if you need a more exhaustive scan.
  4. The test:docker:all command builds the live Docker image once using test:docker:live-build and then reuses it for both live lanes.
  5. Container smoke runners such as test:docker:openwebui, test:docker:onboard, test:docker:gateway-network, test:docker:mcp-channels, and test:docker:plugins boot real containers to verify high-level integration paths.

The live-model Docker runners handle CLI authentication by bind-mounting the necessary homes and copying them into the container. This allows external CLI OAuth to refresh tokens without changing your host’s authentication store.

  1. Direct models:
Terminal window
pnpm test:docker:live-models

(script: scripts/test-live-models-docker.sh)

  1. ACP bind smoke:
Terminal window
pnpm test:docker:live-acp-bind

(script: scripts/test-live-acp-bind-docker.sh)

  1. CLI backend smoke:
Terminal window
pnpm test:docker:live-cli-backend

(script: scripts/test-live-cli-backend-docker.sh)

  1. Codex app-server use smoke:
Terminal window
pnpm test:docker:live-codex-harness

(script: scripts/test-live-codex-use-docker.sh)

  1. Gateway + dev agent:
Terminal window
pnpm test:docker:live-gateway

(script: scripts/test-live-gateway-models-docker.sh)

  1. Open WebUI live smoke:
Terminal window
pnpm test:docker:openwebui

(script: scripts/e2e/openwebui-docker.sh)

  1. Onboarding wizard (TTY, full scaffolding):
Terminal window
pnpm test:docker:onboard

(script: scripts/e2e/onboard-docker.sh)

  1. Gateway networking (two containers, WS auth + health):
Terminal window
pnpm test:docker:gateway-network

(script: scripts/e2e/gateway-network-docker.sh)

  1. MCP channel bridge:
Terminal window
pnpm test:docker:mcp-channels

(script: scripts/e2e/mcp-channels-docker.sh)

  1. Plugins smoke:
Terminal window
pnpm test:docker:plugins

(script: scripts/e2e/plugins-docker.sh)

To keep the runtime image slim, these runners bind-mount your current checkout as read-only and stage it into a temporary directory inside the container. This process skips large caches like .pnpm-store, .worktrees, __openclaw_vitest__, and build outputs so you don’t waste time copying unnecessary files.

  1. The runners set OPENCLAW_SKIP_CHANNELS=1 to prevent starting real Telegram or Discord workers inside the container.
  2. If you use test:docker:openwebui, it starts an OpenClaw Gateway container with OpenAI-compatible API endpoints and verifies the connection through a real chat request.
  3. Successful runs for Open WebUI will print a JSON payload like { "ok": true, "model": "openclaw/default", ... }.
  4. The test:docker:mcp-channels runner is deterministic and doesn’t need real accounts; it verifies the stdio MCP bridge by inspecting raw frames directly.

If you need to perform a manual ACP plain-language thread smoke (not for CI), you can use this script:

Terminal window
bun scripts/dev/discord-acp-plain-language-smoke.ts --channel <discord-channel-id> ...

You should keep this script for debugging workflows, especially for validating ACP thread routing. Here are some useful environment variables you can use:

  1. OPENCLAW_CONFIG_DIR=... (default: ~/.openclaw) mounted to /home/node/.openclaw.
  2. OPENCLAW_WORKSPACE_DIR=... (default: ~/.openclaw/workspace) mounted to /home/node/.openclaw/workspace.
  3. OPENCLAW_PROFILE_FILE=... (default: ~/.profile) mounted to /home/node/.profile and sourced before running tests.
  4. OPENCLAW_DOCKER_PROFILE_ENV_ONLY=1 to verify only env vars sourced from OPENCLAW_PROFILE_FILE, using temporary config/workspace dirs and no external CLI auth mounts.
  5. OPENCLAW_DOCKER_CLI_TOOLS_DIR=... (default: ~/.cache/openclaw/docker-cli-tools) mounted to /home/node/.npm-global for cached CLI installs inside Docker.
  6. External CLI auth dirs/files under $HOME are mounted read-only under /host-auth..., then copied into /home/node/... before tests start. This includes default dirs like .minimax and files like ~/.codex/auth.json, ~/.codex/config.toml, .claude.json, ~/.claude/.credentials.json, ~/.claude/settings.json, and ~/.claude/settings.local.json.
  7. Narrowed provider runs mount only the needed dirs/files inferred from OPENCLAW_LIVE_PROVIDERS / OPENCLAW_LIVE_GATEWAY_PROVIDERS.
  8. Override manually with OPENCLAW_DOCKER_AUTH_DIRS=all, OPENCLAW_DOCKER_AUTH_DIRS=none, or a comma list like OPENCLAW_DOCKER_AUTH_DIRS=.claude,.codex.
  9. OPENCLAW_LIVE_GATEWAY_MODELS=... / OPENCLAW_LIVE_MODELS=... to narrow the run.
  10. OPENCLAW_LIVE_GATEWAY_PROVIDERS=... / OPENCLAW_LIVE_PROVIDERS=... to filter providers in-container.
  11. OPENCLAW_SKIP_DOCKER_BUILD=1 to reuse an existing openclaw:local-live image for reruns that do not need a rebuild.
  12. OPENCLAW_LIVE_REQUIRE_PROFILE_KEYS=1 to ensure creds come from the profile store (not env).
  13. OPENCLAW_OPENWEBUI_MODEL=... to choose the model exposed by the Gateway for the Open WebUI smoke.
  14. OPENCLAW_OPENWEBUI_PROMPT=... to override the nonce-check prompt used by the Open WebUI smoke.
  15. OPENWEBUI_IMAGE=... to override the pinned Open WebUI image tag.

Validate documentation with Docs sanity checks

Section titled “Validate documentation with Docs sanity checks”

Maintaining high-quality documentation is just as important as writing good code. You can use these built-in checks to ensure your edits don’t introduce broken links or formatting issues in the OpenClaw docs.

  1. Run basic documentation checks:
Terminal window
pnpm check:docs
  1. Run full Mintlify anchor validation for in-page heading checks:
Terminal window
pnpm docs:check-links:anchors

You can perform “real pipeline” regressions without needing actual providers by using these CI-safe tests. These tests help you verify that the OpenClaw Gateway and agent loops are functioning as expected in your local environment.

  1. Test the Gateway tool calling by using a mock OpenAI setup alongside the real Gateway and agent loop in src/gateway/gateway.test.ts. You should look for the specific case: “runs a mock OpenAI tool call end-to-end via gateway agent loop”.
  2. Validate the Gateway wizard by checking the WS wizard.start and wizard.next functions in src/gateway/gateway.test.ts. This ensures the system correctly writes configuration files and enforces authentication, specifically in the case: “runs wizard over ws and writes auth token config”.

Improve Agent Reliability and Skills Evaluation

Section titled “Improve Agent Reliability and Skills Evaluation”

You already have access to several CI-safe tests that function as reliability evaluations for your agent. These tests ensure that the core logic of your OpenClaw setup remains stable as you add new features.

  1. Use mock tool-calling through the real Gateway and agent loop as defined in src/gateway/gateway.test.ts.
  2. Run end-to-end wizard flows that validate how sessions are wired together and how configuration changes impact the system in src/gateway/gateway.test.ts.

There are still a few things you need to handle for Skills to be fully tested:

  1. Decisioning: You need to confirm that when skills are listed in a prompt, the agent chooses the correct one and ignores skills that do not apply.
  2. Compliance: You must verify that the agent reads the SKILL.md file before taking action and follows all the required steps and arguments.
  3. Tool Order and History: Assert the correct order of tools in multi-turn scenarios and ensure session history carries over.
  4. Sandbox Boundaries: Confirm that the agent respects the defined boundaries of the sandbox.

Your future evaluations should focus on being deterministic first:

  1. Build a scenario runner using mock providers to assert tool calls, SKILL.md file reads, and session wiring while testing usage, gating, and prompt injection.
  2. Deploy optional live evaluations that are environment-gated only after your CI-safe suite is fully operational.

Master OpenClaw contract tests for plugins and channels

Section titled “Master OpenClaw contract tests for plugins and channels”

OpenClaw contract tests ensure that every plugin and channel you register follows its interface contract perfectly. These tests iterate through all discovered plugins and run a suite of assertions to check both their shape and their behavior.

The default pnpm test unit lane skips these shared files on purpose. You should run these contract commands manually when you modify shared channel or provider surfaces:

  1. Run all contracts using pnpm test:contracts.
  2. Run channel contracts only with pnpm test:contracts:channels.
  3. Run provider contracts only using pnpm test:contracts:plugins.

You can find channel contracts in src/channels/plugins/contracts/*.contract.test.ts. They cover the following areas:

  1. plugin - Basic plugin shape including id, name, capabilities, and metadata.
  2. setup - The setup wizard contract.
  3. session-binding - How sessions are bound.
  4. outbound-payload - The structure of message payloads.
  5. inbound - Handling of inbound messages.
  6. actions - Handlers for channel actions.
  7. threading - How thread IDs are handled.
  8. directory - The directory and roster API.
  9. group-policy - How group policies are enforced.

Provider status contracts are located in src/plugins/contracts/*.contract.test.ts:

  1. status - Probes for channel status.
  2. registry - The shape of the plugin registry.

General provider contracts are also found in src/plugins/contracts/*.contract.test.ts:

  1. auth - The auth flow contract.
  2. auth-choice - Selection of auth choices.
  3. catalog - The model catalog API.
  4. discovery - How plugins are discovered.
  5. loader - How plugins are loaded.
  6. runtime - The provider runtime.
  7. shape - The interface and shape of the plugin.
  8. wizard - The setup wizard.

You should run these tests in these specific scenarios:

  1. After you change plugin-sdk exports or subpaths.
  2. After you add or change a channel or provider plugin.
  3. After you refactor how plugins are registered or discovered.
  4. When you update any shared plugin logic.

These tests run in CI and do not require real API keys to function.

Prevent bugs with OpenClaw regression testing

Section titled “Prevent bugs with OpenClaw regression testing”

When you fix a provider or model issue discovered in a live environment, you should add a regression test to stop it from happening again. This keeps the OpenClaw codebase stable and reliable as you add more features.

  1. Add a CI-safe regression whenever you can. You can use a mock or stub provider, or just capture the exact request-shape transformation to verify the fix.
  2. If the bug only happens in live environments due to things like rate limits or auth policies, keep the test narrow. Make it opt-in using environment variables so it does not impact standard test runs.
  3. Try to hit the smallest layer possible to catch the bug. If it is a provider request conversion or replay bug, use a direct models test. If it is a Gateway session, history, tool pipeline, or state bug, use a Gateway live smoke test or a CI-safe Gateway mock test.
  4. Ensure the regression test is documented within the test suite.

You also need to be aware of the SecretRef traversal guardrail. The test in src/secrets/exec-secret-ref-id-parity.test.ts gets a sampled target for each SecretRef class from registry metadata using listSecretTargetRegistryEntries(). It then checks that traversal-segment exec ids are blocked.

  1. If you add a new includeInPlan SecretRef target family in src/secrets/target-registry-data.ts, you must update classifyTargetClass in that test.
  2. The test will fail on purpose if it sees unclassified target ids so you do not skip new classes by accident.
OpenClaw

OpenClaw Expert

Still stuck?

If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.