Caveat verdict
guardian-wall
guardian-wall-azzar
Prompt injection defense skill that sanitizes external content; behavior is explicitly protective and transparent with no suspicious network usage.
โ Flagged for review โ coarse, uncorroborated signal, not a confirmed exploit. Review the config yourself before installing.
Automated static analysis โ not a human review. Caveat flags capabilities, not confirmed intent, and can produce false positives. Disagree with this verdict? Use Dispute below.
Findings (3)
Prompt injection โ tries to override agent instructions
SKILL.md ยท frontmatter ยท ignore previous instructions
Unicode homoglyph detected โ uses lookalike characters to evade pattern matching
references/patterns.md ยท prose
Possible prompt injection โ attempts to redefine agent identity
scripts/sanitize.py ยท prose ยท downgraded ยท you are now
Permissions & capabilities
No declared permissions โ minimal attack surface.
Is this flag fair?
Thanks โ recorded.