Caveat verdict

llm-regression-monitor

88
🟢 Trusted
No high-risk patterns surfaced by the deep scan — automated capability review, not behavioral proof.

LLM behavioral regression testing system that captures baselines and runs scheduled comparisons; credentials (OPENAI_API_KEY etc.) are used for legitimate LLM API calls to the user's own configured providers.

⚠ Flagged for review — coarse, uncorroborated signal, not a confirmed exploit. Review the config yourself before installing.

Automated static analysis — not a human review. Caveat flags capabilities, not confirmed intent, and can produce false positives. Disagree with this verdict? Use Dispute below.

30
security
70
transparency
70
maintenance

Permission integrity

Installs packages at runtime — transitive dependencies are not auditable

package_install

Findings (6)

Pattern match high

Pipe to python — executes piped content as Python code

SKILL.md · prose · downgraded · | python

Pattern match high

Accesses shell history/config

references/providers.md · code · ~/.zshrc

Pattern match medium

Instructs covert action — may act without user awareness

scripts/send_alert.py · prose · downgraded · silently

Pattern match medium

subprocess execution — runs system commands from Python

scripts/send_alert.py · prose · downgraded · subprocess.run(

Pattern match low

Python os.environ.get — reads environment variable

scripts/capture_baseline.py · prose · downgraded · os.environ.get(

Pattern match low

References webhook/callback URL

scripts/send_alert.py · prose · downgraded · webhook_url

Why the tier is capped

Execution sink present in raw bytes (Hard Floor: class D). Final tier capped at Caution — cannot be lifted by any downgrade, example-payload opt-in, or allowlist.

Permissions & capabilities

No declared permissions — minimal attack surface.

package_installnetwork_in

Is this flag fair?

Check another skill Browse the registry Auditing your own skills or configs? Use the API