Understanding the OpenClaw Threat Model
I have spent many nights worrying about what happens when an AI agent behaves in ways I didn’t expect. Security often feels like a guessing game where you are trying to fix holes you cannot see. It is frustrating to build a great tool only to realize you don’t have a clear map of the risks involved.
I prefer having a structured way to look at threats. That is why I use the OpenClaw Threat Model. It uses the MITRE ATLAS framework to give us a clear view of how to protect our agents and the ClawHub marketplace.
What You’ll Need
Section titled “What You’ll Need”To get started with this model, you should be familiar with these components:
- MITRE ATLAS Framework: The industry standard for AI system threats.
- OpenClaw Agent Runtime: The core where agents execute and call tools.
- Gateway: Handles authentication and routing.
- ClawHub Marketplace: Where skills are published and distributed.
Quick Start
Section titled “Quick Start”You can start using this threat model by focusing on the scope and methodology. I follow these steps to keep my setup secure:
- Identify the Methodology: We use MITRE ATLAS combined with Data Flow Diagrams to track how information moves through the system.
- Check the Scope: Ensure your work covers the Agent Runtime, Gateway, and Channel Integrations (like WhatsApp or Slack).
- Review Marketplace Security: Look at how ClawHub handles skill moderation and distribution.
- Connect External Tools: Include MCP Servers in your security review as they provide external tools to your agents.
Troubleshooting
Section titled “Troubleshooting”If you find a security issue that isn’t covered in the current version, here is how to handle it:
- Unreported Threats: If you find a new risk, follow the guidelines in
CONTRIBUTING-THREAT-MODEL.mdto report it. - Outdated Mitigations: Use the community guidelines to propose updates to existing threat mitigations.
What’s Next
Section titled “What’s Next”If you need help mapping specific attack chains, ask the AI Setup Assistant.
I remember the first time I connected an LLM to a public messaging app. It felt a bit like handing the keys to my house to a stranger who promised to clean up but might accidentally let the cat out—or worse, let a burglar in. When you give an agent the power to use tools and fetch URLs, security can’t be an afterthought.
I want to walk you through how we structured this system. We built it around five specific trust boundaries to ensure that even if one part of the chain is messy, your core system stays safe.
What You’ll Need
Section titled “What You’ll Need”Before we dive into the architecture, make sure you are familiar with these components mentioned in our docs:
- Messaging Channels (WhatsApp, Telegram, or Discord)
- Gateway configuration
- Docker for execution sandboxing
- ClawHub for skill management
Quick Start: The 5-Minute Architecture Tour
Section titled “Quick Start: The 5-Minute Architecture Tour”The system is divided into zones. The “Untrusted Zone” is the wild west of the internet, while the internal layers handle validation, isolation, and execution.
2.1 Trust Boundaries
Section titled “2.1 Trust Boundaries”Here is the high-level map of how data moves from a chat app into your execution environment.
┌─────────────────────────────────────────────────────────────────┐│ UNTRUSTED ZONE ││ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ││ │ WhatsApp │ │ Telegram │ │ Discord │ ... ││ └──────┬──────┘ └──────┬──────┘ └──────┬──────┘ ││ │ │ │ │└─────────┼────────────────┼────────────────┼──────────────────────┘ │ │ │ ▼ ▼ ▼┌─────────────────────────────────────────────────────────────────┐│ TRUST BOUNDARY 1: Channel Access ││ ┌──────────────────────────────────────────────────────────┐ ││ │ GATEWAY │ ││ │ • Device Pairing (30s grace period) │ ││ │ • AllowFrom / AllowList validation │ ││ │ • Token/Password/Tailscale auth │ ││ └──────────────────────────────────────────────────────────┘ │└─────────────────────────────────────────────────────────────────┘ │ ▼┌─────────────────────────────────────────────────────────────────┐│ TRUST BOUNDARY 2: Session Isolation ││ ┌──────────────────────────────────────────────────────────┐ ││ │ AGENT SESSIONS │ ││ │ • Session key = agent:channel:peer │ ││ │ • Tool policies per agent │ ││ │ • Transcript logging │ ││ └──────────────────────────────────────────────────────────┘ │└─────────────────────────────────────────────────────────────────┘ │ ▼┌─────────────────────────────────────────────────────────────────┐│ TRUST BOUNDARY 3: Tool Execution ││ ┌──────────────────────────────────────────────────────────┐ ││ │ EXECUTION SANDBOX │ ││ │ • Docker sandbox OR Host (exec-approvals) │ ││ │ • Node remote execution │ ││ │ • SSRF protection (DNS pinning + IP blocking) │ ││ └──────────────────────────────────────────────────────────┘ │└─────────────────────────────────────────────────────────────────┘ │ ▼┌─────────────────────────────────────────────────────────────────┐│ TRUST BOUNDARY 4: External Content ││ ┌──────────────────────────────────────────────────────────┐ ││ │ FETCHED URLs / EMAILS / WEBHOOKS │ ││ │ • External content wrapping (XML tags) │ ││ │ • Security notice injection │ ││ └──────────────────────────────────────────────────────────┘ │└─────────────────────────────────────────────────────────────────┘ │ ▼┌─────────────────────────────────────────────────────────────────┐│ TRUST BOUNDARY 5: Supply Chain ││ ┌──────────────────────────────────────────────────────────┐ ││ │ CLAWHUB │ ││ │ • Skill publishing (semver, SKILL.md required) │ ││ │ • Pattern-based moderation flags │ ││ │ • VirusTotal scanning (coming soon) │ ││ │ • GitHub account age verification │ ││ └──────────────────────────────────────────────────────────┘ │└─────────────────────────────────────────────────────────────────┘1. Channel Access (The Gateway)
Section titled “1. Channel Access (The Gateway)”This is your first line of defense. When you pair a device, there is a strict 30s grace period. I recommend using AllowFrom or AllowList validation to ensure only specific users can talk to your agent. You can also use Tailscale or standard Token/Password auth here.
2. Session Isolation
Section titled “2. Session Isolation”Every interaction is siloed. We use a session key format of agent:channel:peer. This ensures that data from one user on Telegram doesn’t leak into a session with another user on Discord. We also apply tool policies here and keep full transcript logs.
3. Tool Execution
Section titled “3. Tool Execution”This is where the heavy lifting happens. You can run tools in a Docker sandbox or directly on the Host if you use exec-approvals. To prevent the agent from attacking your internal network, we use SSRF protection including DNS pinning and IP blocking.
4. External Content
Section titled “4. External Content”When the agent fetches a URL or reads an email, that content is untrusted. We wrap this external data in XML tags and inject security notices so the LLM knows it is looking at outside information.
5. Supply Chain (ClawHub)
Section titled “5. Supply Chain (ClawHub)”If you are using third-party skills, ClawHub handles the vetting. It requires semver and a SKILL.md file. It also checks GitHub account age and will soon include VirusTotal scanning.
2.2 Data Flows
Section titled “2.2 Data Flows”To help you visualize how data moves through these boundaries, I’ve included this flow table from the source docs:
| Flow | Source | Destination | Data | Protection |
|---|---|---|---|---|
| F1 | Channel | Gateway | User messages | TLS, AllowFrom |
| F2 | Gateway | Agent | Routed messages | Session isolation |
| F3 | Agent | Tools | Tool invocations | Policy enforcement |
| F4 | Agent | External | web_fetch requests | SSRF blocking |
| F5 | ClawHub | Agent | Skill code | Moderation, scanning |
| F6 | Agent | Channel | Responses | Output filtering |
Troubleshooting
Section titled “Troubleshooting”If things aren’t working as expected, it’s usually a boundary issue. Here are two common problems:
- Pairing Fails: If you can’t connect your device, check the clock. The Gateway has a 30s grace period for device pairing. If you miss that window, the request is rejected.
- Skill Upload Rejected: ClawHub has strict requirements. Make sure your repository has a
SKILL.mdfile and follows semver. Also, if your GitHub account is brand new, the age verification might flag it. - Tool Request Blocked: If a tool fails to fetch a URL, the SSRF protection might be active. Check if you are hitting a blocked IP range or if DNS pinning is interfering with a local dev setup.
- Unauthorized Access: Double-check your
AllowListorAllowFromsettings in the Gateway. If these aren’t configured, the gateway might be blocking incoming messages from unknown peers.
If you have more questions about setting this up, check out the AI Setup Assistant.
What’s Next
Section titled “What’s Next”- Gateway Configuration Details
- Setting up Docker Sandboxes
- Publishing to ClawHub
- Managing Agent Sessions
I used to think that securing an AI agent was just about writing a good system prompt. Then I realized that if someone can trick my bot into reading a malicious URL, they might be able to execute commands on my host machine. It’s a bit of a wake-up call when you realize your helpful assistant could be turned into a proxy for an attacker.
I want to walk you through how we analyze these threats using the ATLAS framework. It’s not just about “hacking”; it’s about understanding the specific ways an LLM-based system can be compromised, from initial access to full system impact.
What You’ll Need
Section titled “What You’ll Need”- OpenClaw Gateway access
- A configured
~/.openclaw/credentials/directory - Active channel integrations (like messaging platforms)
- Access to the
exec-approvals.tsconfiguration
Quick Start: Secure Your Agent in 5 Minutes
Section titled “Quick Start: Secure Your Agent in 5 Minutes”If you want to minimize your risk immediately, I recommend these four steps based on our current mitigations:
- Network Lockdown: Ensure your gateway is bound to
loopbackby default or use the Tailscale auth option to prevent public discovery. - Enable Sanboxing: Use the Docker sandbox option for the Bash tool to prevent unauthorized command execution on your host.
- Strict Approvals: Keep “ask mode” enabled in
exec-approvals.tsfor all dangerous commands. - Token Safety: Check your file permissions on
~/.openclaw/credentials/to ensure other users on your system can’t read your tokens.
Threat Analysis by ATLAS Tactic
Section titled “Threat Analysis by ATLAS Tactic”I’ve broken down the risks we’re tracking. Each one follows the ATLAS Tactic IDs so you can map them to industry-standard AI security research.
1. Reconnaissance (AML.TA0002)
Section titled “1. Reconnaissance (AML.TA0002)”Attackers start by looking for a way in.
- T-RECON-001 (Agent Endpoint Discovery): Attackers might use Shodan queries or DNS enumeration to find exposed OpenClaw gateways. We currently mitigate this by binding to loopback by default and offering Tailscale authentication.
- T-RECON-002 (Channel Integration Probing): Someone might send test messages to your messaging accounts to see if an AI responds. This helps them identify which accounts are AI-managed.
2. Initial Access (AML.TA0004)
Section titled “2. Initial Access (AML.TA0004)”How does an attacker get control?
- T-ACCESS-001 (Pairing Code Interception): During the 30-second grace period when you pair a device, a local attacker could try to intercept that code.
- T-ACCESS-002 (AllowFrom Spoofing): Depending on the channel, an attacker might spoof a phone number or username to bypass your
AllowFromfilters. - T-ACCESS-003 (Token Theft): If an attacker gets malware on your machine, they could steal plaintext tokens from
~/.openclaw/credentials/. We rely on OS-level file permissions to stop this right now.
3. Execution (AML.TA0005)
Section titled “3. Execution (AML.TA0005)”This is where the LLM itself is targeted.
- T-EXEC-001 (Direct Prompt Injection): An attacker sends a message directly to the agent to manipulate its behavior. We use pattern detection and external content wrapping to catch these.
- T-EXEC-002 (Indirect Prompt Injection): This is sneakier. An attacker hides instructions in a website that your agent fetches via
web_fetch. We wrap this content in XML tags with security notices to tell the LLM it’s untrusted data. - T-EXEC-003 (Tool Argument Injection): An attacker tries to influence the parameters sent to a tool. For example, tricking the agent into deleting a file by manipulating a “filename” argument.
- T-EXEC-004 (Exec Approval Bypass): Attackers might use command aliases or path manipulation to bypass the allowlist in
exec-approvals.ts.
4. Persistence (AML.TA0006)
Section titled “4. Persistence (AML.TA0006)”How they stay in your system.
- T-PERSIST-001 (Malicious Skill Installation): An attacker could publish a malicious skill to ClawHub. We check GitHub account age and use pattern-based moderation to flag these.
- T-PERSIST-002 (Skill Update Poisoning): If a popular skill is compromised, an auto-update could pull in malicious code. We currently use version fingerprinting to track changes.
5. Collection & Exfiltration (AML.TA0009, AML.TA0010)
Section titled “5. Collection & Exfiltration (AML.TA0009, AML.TA0010)”The goal is often your data.
- T-EXFIL-001 (Data Theft via web_fetch): An attacker might use prompt injection to make your agent
POSTsensitive data to their own external URL. We block SSRF for internal networks, but external URLs are still a risk. - T-EXFIL-002 (Unauthorized Message Sending): The agent could be tricked into messaging sensitive info to the attacker’s account.
- T-EXFIL-003 (Credential Harvesting): A malicious skill could try to read your environment variables or config files. Since skills run with agent privileges, this is a high-risk area.
6. Impact (AML.TA0011)
Section titled “6. Impact (AML.TA0011)”The final result of an attack.
- T-IMPACT-001 (Unauthorized Command Execution): The worst-case scenario where an attacker runs arbitrary commands on your system.
- T-IMPACT-002 (Resource Exhaustion): An attacker could flood your agent with messages to exhaust your API credits or compute power.
- T-IMPACT-003 (Reputation Damage): Tricking the agent into sending offensive content to others.
Troubleshooting
Section titled “Troubleshooting”Here are a few specific issues you might run into based on our threat model:
Your gateway is showing up in public scans
- Check your binding settings. If you aren’t using a VPN like Tailscale, ensure the gateway isn’t listening on
0.0.0.0.
The agent is ignoring security wrappers in fetched content
- This happens with “Indirect Prompt Injection” (T-EXEC-002). If the LLM ignores the XML tags, you may need to use a more advanced model or implement manual output validation.
A skill is requesting weird permissions
- This could be a sign of T-EXFIL-003. Check the skill source code on ClawHub and verify the GitHub account of the author.
If you have questions about a specific ATLAS ID or need help hardening your setup, check out the AI Setup Assistant.
What’s Next
Section titled “What’s Next”- Configuring exec-approvals.ts
- Setting up the Docker Sandbox
- ClawHub Security Policies
- Using Tailscale with OpenClaw
I’ve often felt that uneasy chill when pulling in external code for a project. You want the functionality, but you don’t want the baggage of a security breach. It’s a common stress for anyone building platforms where users share their work. I want to share how we are currently looking at these supply chain risks and what we are doing to keep things safe.
What You’ll Need
Section titled “What You’ll Need”To understand the current security posture, I’m looking at these specific controls:
- GitHub account age verification via
requireGitHubAccountAge() - Path sanitization using
sanitizePath() - File type validation via
isTextFile() - Total bundle size limits (50MB)
Quick Start
Section titled “Quick Start”If you are looking to understand the 5-minute path of our current security flow, here is how the system handles a new skill submission:
- Account Check: The system runs
requireGitHubAccountAge()to ensure the user isn’t using a brand-new throwaway account. - Path Scrubbing: Every file path goes through
sanitizePath()to stop path traversal attacks dead in their tracks. - Type & Size Check: We use
isTextFile()to ensure only text is uploaded and verify the total bundle stays under 50MB. - Pattern Scan: The system checks the slug and metadata against our
FLAG_RULESinmoderation.ts.
Current Security Controls
Section titled “Current Security Controls”I’ve broken down the effectiveness of what we have running right now. Some things work great, while others are just a starting point.
| Control | Implementation | Effectiveness |
|---|---|---|
| GitHub Account Age | requireGitHubAccountAge() | Medium - Raises bar for new attackers |
| Path Sanitization | sanitizePath() | High - Prevents path traversal |
| File Type Validation | isTextFile() | Medium - Only text files, but can still be malicious |
| Size Limits | 50MB total bundle | High - Prevents resource exhaustion |
| Required SKILL.md | Mandatory readme | Low security value - Informational only |
| Pattern Moderation | FLAG_RULES in moderation.ts | Low - Easily bypassed |
| Moderation Status | moderationStatus field | Medium - Manual review possible |
Moderation Flag Patterns
Section titled “Moderation Flag Patterns”We use specific regex patterns in moderation.ts to catch obvious bad actors. Here is the actual code the system uses to scan slugs, display names, and metadata:
// Known-bad identifiers/(keepcold131\/ClawdAuthenticatorTool|ClawdAuthenticatorTool)/i
// Suspicious keywords/(malware|stealer|phish|phishing|keylogger)/i/(api[-_ ]?key|token|password|private key|secret)/i/(wallet|seed phrase|mnemonic|crypto)/i/(discord\.gg|webhook|hooks\.slack)/i/(curl[^\n]+\|\s*(sh|bash))/i/(bit\.ly|tinyurl\.com|t\.co|goo\.gl|is\.gd)/iI should point out that this has limitations. It only checks the surface-level metadata and doesn’t look at the actual skill code. Simple obfuscation can bypass these regex rules.
Risk Matrix
Section titled “Risk Matrix”I use a risk matrix to prioritize what we fix first. We look at likelihood versus impact to determine the risk level.
| Threat ID | Likelihood | Impact | Risk Level | Priority |
|---|---|---|---|---|
| T-EXEC-001 | High | Critical | Critical | P0 |
| T-PERSIST-001 | High | Critical | Critical | P0 |
| T-EXFIL-003 | Medium | Critical | Critical | P0 |
| T-IMPACT-001 | Medium | Critical | High | P1 |
| T-EXEC-002 | High | High | High | P1 |
| T-EXEC-004 | Medium | High | High | P1 |
| T-ACCESS-003 | Medium | High | High | P1 |
| T-EXFIL-001 | Medium | High | High | P1 |
| T-IMPACT-002 | High | Medium | High | P1 |
| T-EVADE-001 | High | Medium | Medium | P2 |
Critical Path Attack Chains
Section titled “Critical Path Attack Chains”These are the sequences I’m most worried about right now:
- Skill-Based Data Theft:
T-PERSIST-001→T-EVADE-001→T-EXFIL-003. This happens when someone publishes a malicious skill, evades moderation, and harvests credentials. - Prompt Injection to RCE:
T-EXEC-001→T-EXEC-004→T-IMPACT-001. This involves injecting a prompt to bypass execution approval and running commands. - Indirect Injection:
T-EXEC-002→T-EXFIL-001→ External exfiltration. This is where an attacker poisons URL content that an agent fetches, leading to data being sent out.
Recommendations Summary
Section titled “Recommendations Summary”I’ve categorized our next steps into three priority levels.
Immediate (P0)
Section titled “Immediate (P0)”- R-001: Complete VirusTotal integration to address
T-PERSIST-001. - R-002: Implement skill sandboxing to stop data exfiltration.
- R-003: Add output validation for sensitive actions.
Short-term (P1)
Section titled “Short-term (P1)”- R-004: Implement rate limiting.
- R-005: Add token encryption at rest.
- R-006: Improve exec approval UX and validation.
- R-007: Implement URL allowlisting for
web_fetch.
Medium-term (P2)
Section titled “Medium-term (P2)”- R-008: Add cryptographic channel verification.
- R-009: Implement config integrity verification.
- R-010: Add update signing and version pinning.
Troubleshooting
Section titled “Troubleshooting”If you’re seeing issues with the current moderation system, here are two common areas where things go wrong:
- Bypassed Flags: If a malicious skill gets through, it’s likely because the attacker used obfuscation. We are moving toward VirusTotal behavioral analysis to fix this.
- Manual Review Delays: If a skill is stuck in “Manual Review,” it’s because it triggered a
moderationStatusflag. You can check theauditLogstable to see why it was flagged.
If you need more help setting up your security environment, check out the AI Setup Assistant.
What’s Next
Section titled “What’s Next”- Check out the
skillReportstable documentation. - Review the
auditLogsimplementation details. - Read about our Planned Improvements.
- See the full Risk Matrix.
I’ve often found myself staring at security reports wondering which part of my code is actually responsible for stopping a specific attack. It’s frustrating when documentation is vague about where the “scary stuff” is handled. I want to walk you through the reference tables we use to keep OpenClaw safe, so you know exactly which files to watch and how we map to industry standards.
What You’ll Need
Section titled “What You’ll Need”- Access to the OpenClaw source code repository
- A basic understanding of the MITRE ATLAS framework
Quick Start
Section titled “Quick Start”I recommend using these tables to audit your local setup. If you are looking for specific logic—like how we handle SSRF or prompt injection—start with the file mapping.
1. Identify the Threat (ATLAS Mapping)
Section titled “1. Identify the Threat (ATLAS Mapping)”We map our internal threat IDs to the MITRE ATLAS (Adversarial Threat Landscape for AI Systems) framework. This helps you understand the intent behind our security controls.
| ATLAS ID | Technique Name | OpenClaw Threats |
|---|---|---|
| AML.T0006 | Active Scanning | T-RECON-001, T-RECON-002 |
| AML.T0009 | Collection | T-EXFIL-001, T-EXFIL-002, T-EXFIL-003 |
| AML.T0010.001 | Supply Chain: AI Software | T-PERSIST-001, T-PERSIST-002 |
| AML.T0010.002 | Supply Chain: Data | T-PERSIST-003 |
| AML.T0031 | Erode AI Model Integrity | T-IMPACT-001, T-IMPACT-002, T-IMPACT-003 |
| AML.T0040 | AI Model Inference API Access | T-ACCESS-001, T-ACCESS-002, T-ACCESS-003, T-DISC-001, T-DISC-002 |
| AML.T0043 | Craft Adversarial Data | T-EXEC-004, T-EVADE-001, T-EVADE-002 |
| AML.T0051.000 | LLM Prompt Injection: Direct | T-EXEC-001, T-EXEC-003 |
| AML.T0051.001 | LLM Prompt Injection: Indirect | T-EXEC-002 |
2. Locate the Code (Key Security Files)
Section titled “2. Locate the Code (Key Security Files)”If you need to modify or audit security logic, these are the files you should look at. I’ve categorized them by risk level so you know which ones are the most sensitive.
| Path | Purpose | Risk Level |
|---|---|---|
src/infra/exec-approvals.ts | Command approval logic | Critical |
src/gateway/auth.ts | Gateway authentication | Critical |
src/web/inbound/access-control.ts | Channel access control | Critical |
src/infra/net/ssrf.ts | SSRF protection | Critical |
src/security/external-content.ts | Prompt injection mitigation | Critical |
src/agents/sandbox/tool-policy.ts | Tool policy enforcement | Critical |
convex/lib/moderation.ts | ClawHub moderation | High |
convex/lib/skillPublish.ts | Skill publishing flow | High |
src/routing/resolve-route.ts | Session isolation | Medium |
Troubleshooting
Section titled “Troubleshooting”If you encounter terms in our threat model or codebase that seem unclear, use this glossary to get back on track.
- ATLAS: MITRE’s Adversarial Threat Landscape for AI Systems.
- ClawHub: OpenClaw’s skill marketplace.
- Gateway: OpenClaw’s message routing and authentication layer.
- MCP: Model Context Protocol - tool provider interface.
- Prompt Injection: Attack where malicious instructions are embedded in input.
- Skill: Downloadable extension for OpenClaw agents.
- SSRF: Server-Side Request Forgery.
This threat model is a living document. If you find a security issue, please report it to security@openclaw.ai.
Still have questions about these mappings or specific files? Ask our AI Setup Assistant.
What’s Next
Section titled “What’s Next”OpenClaw Expert
Still stuck?
If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.