Caveat verdict

agent-hardening

agent-hardening-zurbrick

88
🟢 Trusted
No high-risk patterns surfaced by the deep scan — automated capability review, not behavioral proof.

A security hardening guide for LLM agents covering prompt injection defense, MCP permission auditing, and behavioral rules; the content is purely instructional and defensive in nature with no exfiltration or malicious behavior.

⚠ Flagged for review — coarse, uncorroborated signal, not a confirmed exploit. Review the config yourself before installing.

Automated static analysis — not a human review. Caveat flags capabilities, not confirmed intent, and can produce false positives. Disagree with this verdict? Use Dispute below.

35
security
70
transparency
70
maintenance

Findings (6)

Pattern match high

Prompt injection — tries to override agent instructions

references/quick-test.md · prose · downgraded · Ignore all previous instructions

Pattern match high

Accesses sensitive system files

references/quick-test.md · prose · downgraded · /etc/passwd

Pattern match medium

Possible prompt injection — attempts to redefine agent identity

references/quick-test.md · prose · downgraded · You are now

Pattern match medium

Fake system prompt — attempts to inject instructions

references/quick-test.md · prose · downgraded · System: You are

Pattern match low

Python urllib.request — network access

references/quick-test.md · prose · downgraded · urllib.request

Pattern match low

Python os.environ.get — reads environment variable

tools/run-security-tests.py · prose · downgraded · os.environ.get(

Why the tier is capped

Execution sink present in raw bytes (Hard Floor: class C). Final tier capped at Caution — cannot be lifted by any downgrade, example-payload opt-in, or allowlist.

Permissions & capabilities

No declared permissions — minimal attack surface.

Is this flag fair?

Check another skill Browse the registry Auditing your own skills or configs? Use the API