Caveat verdict
agent-hardening
agent-hardening-zurbrick
A security hardening guide for LLM agents covering prompt injection defense, MCP permission auditing, and behavioral rules; the content is purely instructional and defensive in nature with no exfiltration or malicious behavior.
⚠ Flagged for review — coarse, uncorroborated signal, not a confirmed exploit. Review the config yourself before installing.
Automated static analysis — not a human review. Caveat flags capabilities, not confirmed intent, and can produce false positives. Disagree with this verdict? Use Dispute below.
Findings (6)
Prompt injection — tries to override agent instructions
references/quick-test.md · prose · downgraded · Ignore all previous instructions
Accesses sensitive system files
references/quick-test.md · prose · downgraded · /etc/passwd
Possible prompt injection — attempts to redefine agent identity
references/quick-test.md · prose · downgraded · You are now
Fake system prompt — attempts to inject instructions
references/quick-test.md · prose · downgraded · System: You are
Python urllib.request — network access
references/quick-test.md · prose · downgraded · urllib.request
Python os.environ.get — reads environment variable
tools/run-security-tests.py · prose · downgraded · os.environ.get(
Why the tier is capped
Execution sink present in raw bytes (Hard Floor: class C). Final tier capped at Caution — cannot be lifted by any downgrade, example-payload opt-in, or allowlist.
Permissions & capabilities
No declared permissions — minimal attack surface.
Is this flag fair?
Thanks — recorded.