Back to skill

Security audit

self-improving agent

Security checks for vulnerabilities and agentic risk

Overview

The skill is a disclosed local learning/memory workflow with optional hooks; it is not malicious, but users should understand it persists notes and can update workspace guidance files.

Install this only if you want OpenClaw to keep local persistent learning notes and, when appropriate, update workspace guidance files. Keep .learnings out of version control unless you intentionally want to share it, review promoted SOUL.md/TOOLS.md/AGENTS.md changes, and enable the hook only if you accept session-end transcript scanning for redacted error excerpts.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (24)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared purpose is note-taking and review of learnings, but the skill also includes installation, hook enablement, cross-session communication, promotion into workspace prompt files, and skill extraction workflows. This mismatch can cause agents or users to trust and auto-activate the skill in low-risk contexts while it actually performs broader state-changing operations across the workspace.

Content

No source excerpt is available for this finding.

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · SKILL.md (reported line 199)May include surrounding context.

md
4. **Choose** `retain` (supported and useful, or explicitly provisional while blocked), `revise` (correct/narrow with evidence), `externalize` (replace volatile detail with a usable retrieval instruction), or `retire` (exclude from active guidance with an evidenced reason). Externalization names where to look, what to read/test, and what to do if inaccessible/conflicting; a vague “check docs” is insufficient. Verify that the retrieval path is usable before treating it as validated.
5. **Record and propagate**: append a dated review note with the outcome, reason, evidence, and old → new claim when changed. Preserve original observations, resolution, promotion targets/dates, and prior reviews. Update affected promoted guidance using its source link; keep its source-learning ID beside the revised rule (or a retirement note pointing to the preserved history). Never inflate `Recurrence-Count`/`Last-Seen` just because a review ran.

Never erase learning/promotion history or delete by age. Never change user preferences, weaken security policies, or rewrite skill assets merely because they are old or unused. Skill creation/modification requires explicit scoped approval; approval for this integration does not authorize future skill rewrites. If a needed target edit is outside granted scope, record the proposed change as pending in the source entry and report the still-active guidance rather than silently claiming it is updated.

For worked review outcomes and provenance, read [maintenance examples](references/examples.md#maintenance-examples).

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · assets/LEARNINGS.md (reported line 78)May include surrounding context.

Skill-Path: skills/skill-name

text

Only create/modify skill assets with explicit scoped approval. Record the promotion date and target section in the learning's history/Guidance, and the original file + learning ID in the skill. `promoted_to_skill` does not establish current validity.

Example:
```markdown

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 204)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/uninstall.md (reported line 24)May include surrounding context.

learnings directory instead (archive it first if it has entries):

bash
rm -r ~/.openclaw/workspace/.learnings

Restart the gateway after hook changes.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/uninstall.md (reported line 34)May include surrounding context.

learnings directory instead (archive it first if it has entries):

bash
rm -r ~/.openclaw/workspace/.learnings

Restart the gateway after hook changes.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/uninstall.md (reported line 37)May include surrounding context.

learnings directory instead (archive it first if it has entries):

bash
rm -r ~/.openclaw/workspace/.learnings

Restart the gateway after hook changes.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/uninstall.md (reported line 40)May include surrounding context.

learnings directory instead (archive it first if it has entries):

bash
rm -r ~/.openclaw/workspace/.learnings

Restart the gateway after hook changes.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/uninstall.md (reported line 24)May include surrounding context.

learnings directory instead (archive it first if it has entries):

bash
rm -r ~/.openclaw/workspace/.learnings

Restart the gateway after hook changes.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/uninstall.md (reported line 40)May include surrounding context.

learnings directory instead (archive it first if it has entries):

bash
rm -r ~/.openclaw/workspace/.learnings

Restart the gateway after hook changes.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/uninstall.md (reported line 34)May include surrounding context.

bash
# 1. Disable and remove the hook
openclaw hooks disable self-improvement
rm -r ~/.openclaw/hooks/self-improvement

# 2. Remove the skill
rm -r ~/.openclaw/skills/self-improving-agent

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/uninstall.md (reported line 37)May include surrounding context.

md
rm -r ~/.openclaw/hooks/self-improvement

# 2. Remove the skill
rm -r ~/.openclaw/skills/self-improving-agent

# 3. Optional — remove captured learnings (REVIEW FIRST, this is your data)
rm -r ~/.openclaw/workspace/.learnings

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill instructs use of shell commands, git clone/cp, OpenClaw installers, optional hooks, and cross-session features, but it does not declare any explicit tool or permission scope. That means a host agent may invoke filesystem, environment, or network-capable actions without a machine-readable boundary, increasing the chance of over-privileged execution or unsafe auto-enablement.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The activation description covers many normal situations such as failed commands, user corrections, missing capabilities, outdated knowledge, and better approaches. Because these are common during ordinary work, the skill may trigger frequently and opportunistically, causing persistent writes and workflow changes in contexts where the user did not intend memory capture or review behavior.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
84% confidence
Finding

The skill persists data in ~/.openclaw/workspace/.learnings, making captured information survive the current session and potentially affect future sessions. Persistent storage is part of the stated purpose, but it becomes risky because the same document also encourages broad automatic capture and later promotion into injected workspace guidance files.

Content

Scanner excerpt · SKILL.md (reported line 94)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The 'Automatically log when you notice' section directs autonomous capture of corrections, feature requests, knowledge gaps, and errors based on broad signals. In practice this creates implicit persistence of conversation-derived data and can over-collect sensitive operational context, especially combined with review/promotion flows and optional hooks.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The feature-request trigger phrases are broad conversational patterns like 'Can you also...' and 'Is there a way to...', which commonly occur in normal dialogue and do not necessarily indicate consent to persistent logging. This can lead to silent retention of user intent or project details in .learnings/FEATURE_REQUESTS.md without a specific opt-in boundary.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · CHANGELOG.md (reported line 180)May include surrounding context.

Enable the error sweep by creating the learnings directory:

bash
mkdir -p ~/.openclaw/workspace/.learnings

Testing

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · hooks/openclaw/HOOK.md (reported line 71)May include surrounding context.

Enable the error sweep by creating the learnings directory:

bash
mkdir -p ~/.openclaw/workspace/.learnings

Testing

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 266)May include surrounding context.

Enable the error sweep by creating the learnings directory:

bash
mkdir -p ~/.openclaw/workspace/.learnings

Testing

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

Historical extraction example, not current platform advice. Before extracting or reusing it, check the exact image manifest, target platform, runtime/configuration, and representative build/run behavior. Do not generalize one image's missing ARM variant to all Apple Silicon builds. Creating or changing the skill also needs scoped approval.

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 75)May include surrounding context.

md
`<workspace>/.learnings/ERRORS.md` (only if `.learnings/` exists — see
  [Error Detection](#error-detection))

### 3. Create Learning Files

Create the `.learnings/` directory in your workspace:

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This markdown file explicitly defines detection triggers, and several listed triggers such as 'Knowledge gaps' and user corrections like 'No, that's wrong...' overlap with common conversational situations. The section does not provide scope limits, exclusions, or negative examples to clarify when these triggers should or should not apply.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/uninstall.md (reported line 11)May include surrounding context.

md
> errors, corrections, and insights the skill captured for you. Review or
> archive it before deleting anything. Content the skill *promoted* into
> `SOUL.md`, `TOOLS.md`, or `AGENTS.md` is part of those files now — removing
> the skill does not (and should not) automatically remove it.

## Disable only

Static analysis

No suspicious patterns detected.