Caveat verdict
llm-regression-monitor
LLM behavioral regression testing system that captures baselines and runs scheduled comparisons; credentials (OPENAI_API_KEY etc.) are used for legitimate LLM API calls to the user's own configured providers.
⚠ Flagged for review — coarse, uncorroborated signal, not a confirmed exploit. Review the config yourself before installing.
Automated static analysis — not a human review. Caveat flags capabilities, not confirmed intent, and can produce false positives. Disagree with this verdict? Use Dispute below.
Permission integrity
package_install
Findings (6)
Pipe to python — executes piped content as Python code
SKILL.md · prose · downgraded · | python
Accesses shell history/config
references/providers.md · code · ~/.zshrc
Instructs covert action — may act without user awareness
scripts/send_alert.py · prose · downgraded · silently
subprocess execution — runs system commands from Python
scripts/send_alert.py · prose · downgraded · subprocess.run(
Python os.environ.get — reads environment variable
scripts/capture_baseline.py · prose · downgraded · os.environ.get(
References webhook/callback URL
scripts/send_alert.py · prose · downgraded · webhook_url
Why the tier is capped
Execution sink present in raw bytes (Hard Floor: class D). Final tier capped at Caution — cannot be lifted by any downgrade, example-payload opt-in, or allowlist.
Permissions & capabilities
No declared permissions — minimal attack surface.
package_installnetwork_in Is this flag fair?
Thanks — recorded.