Caveat verdict

indirect-prompt-injection

88
๐ŸŸข Trusted
No high-risk patterns surfaced by the deep scan โ€” automated capability review, not behavioral proof.

A defensive security skill that teaches detection and rejection of prompt injection attacks in external content; the skill describes attack patterns as examples for detection, not as payloads to execute.

โš  Flagged for review โ€” coarse, uncorroborated signal, not a confirmed exploit. Review the config yourself before installing.

Automated static analysis โ€” not a human review. Caveat flags capabilities, not confirmed intent, and can produce false positives. Disagree with this verdict? Use Dispute below.

0
security
90
transparency
70
maintenance

Findings (4)

Pattern match critical

Prompt injection โ€” tries to override agent instructions

references/attack-patterns.md ยท code ยท Ignore all previous instructions

Pattern match critical

Unicode homoglyph detected โ€” uses lookalike characters to evade pattern matching

tests/test_cases.json ยท prose

Pattern match high

Possible prompt injection โ€” attempts to redefine agent identity

references/attack-patterns.md ยท code ยท You are now

Pattern match high

Accesses .ssh directory

references/attack-patterns.md ยท code ยท .ssh/

Permissions & capabilities

No declared permissions โ€” minimal attack surface.

Is this flag fair?

Check another skill Browse the registry Auditing your own skills or configs? Use the API