Caveat verdict
indirect-prompt-injection
A defensive security skill that teaches detection and rejection of prompt injection attacks in external content; the skill describes attack patterns as examples for detection, not as payloads to execute.
โ Flagged for review โ coarse, uncorroborated signal, not a confirmed exploit. Review the config yourself before installing.
Automated static analysis โ not a human review. Caveat flags capabilities, not confirmed intent, and can produce false positives. Disagree with this verdict? Use Dispute below.
Findings (4)
Prompt injection โ tries to override agent instructions
references/attack-patterns.md ยท code ยท Ignore all previous instructions
Unicode homoglyph detected โ uses lookalike characters to evade pattern matching
tests/test_cases.json ยท prose
Possible prompt injection โ attempts to redefine agent identity
references/attack-patterns.md ยท code ยท You are now
Accesses .ssh directory
references/attack-patterns.md ยท code ยท .ssh/
Permissions & capabilities
No declared permissions โ minimal attack surface.
Is this flag fair?
Thanks โ recorded.