Caveat verdict

eval-driven-dev

88
๐ŸŸข Trusted
No high-risk patterns surfaced by the deep scan โ€” automated capability review, not behavioral proof.

Receives external input AND uses eval

Evaluation-driven development skill for Python LLM applications using the pixie-qa framework; dynamic_eval capability reflects running Python test code in the development workflow, not arbitrary untrusted code execution.

Automated static analysis โ€” not a human review. Caveat flags capabilities, not confirmed intent, and can produce false positives. Disagree with this verdict? Use Dispute below.

43
security
50
transparency
70
maintenance

What it does

These are capability combinations: each listed behavior occurs in the skill, but Caveat detects co-occurrence โ€” it does not verify that one flows into another. Read the code to confirm a live chain.

Capability combination critical

Receives external input AND uses eval โ€” the remote code-injection pattern (data-flow not verified)

LLM01 ยท LLM05 ยท ASI01 ยท ASI05

Permission integrity

Installs packages at runtime โ€” transitive dependencies are not auditable

package_install

Accesses agent memory/configuration files

agent_memory

Findings (2)

Pattern match medium

References agent memory files

SKILL.md ยท code ยท MEMORY.md

Pattern match low

Python urllib.request โ€” network access

resources/check_version.py ยท prose ยท downgraded ยท urllib.request

Why the tier is capped

Execution sink present in raw bytes (Hard Floor: class B). Final tier capped at Caution โ€” cannot be lifted by any downgrade, example-payload opt-in, or allowlist.

Permissions & capabilities

No declared permissions โ€” minimal attack surface.

dynamic_evalagent_memorypackage_installcredential_accessnetwork_in
Check another skill Browse the registry Auditing your own skills or configs? Use the API