Caveat verdict

agent-benchmark

45
๐ŸŸ  Risky
Significant risk patterns flagged โ€” automated deep scan, not behavioral proof.

The skill runs benchmark tests using PowerShell, which could potentially be used for malicious purposes if not properly validated, but seems to be for performance evaluation.

โš  Flagged for review โ€” coarse, uncorroborated signal, not a confirmed exploit. Review the config yourself before installing.

Automated static analysis โ€” not a human review. Caveat flags capabilities, not confirmed intent, and can produce false positives. Disagree with this verdict? Use Dispute below.

78
security
80
transparency
80
maintenance

Findings (3)

Pattern match medium

References child_process โ€” can spawn system processes

index.js ยท prose ยท downgraded ยท child_process

Pattern match medium

Uses spawn() โ€” can execute external programs

index.js ยท prose ยท downgraded ยท spawn(

Pattern match low

Popular HTTP library โ€” network access

TEST_SUMMARY.md ยท prose ยท downgraded ยท got

Why the tier is capped

Execution sink present in raw bytes (Hard Floor: class D/E). Final tier capped at Caution โ€” cannot be lifted by any downgrade, example-payload opt-in, or allowlist.

Permissions & capabilities

No declared permissions โ€” minimal attack surface.

Is this flag fair?

Check another skill Browse the registry Auditing your own skills or configs? Use the API