Caveat verdict
eval-driven-dev
Receives external input AND uses eval
Evaluation-driven development skill for Python LLM applications using the pixie-qa framework; dynamic_eval capability reflects running Python test code in the development workflow, not arbitrary untrusted code execution.
Automated static analysis โ not a human review. Caveat flags capabilities, not confirmed intent, and can produce false positives. Disagree with this verdict? Use Dispute below.
What it does
These are capability combinations: each listed behavior occurs in the skill, but Caveat detects co-occurrence โ it does not verify that one flows into another. Read the code to confirm a live chain.
Receives external input AND uses eval โ the remote code-injection pattern (data-flow not verified)
LLM01 ยท LLM05 ยท ASI01 ยท ASI05
Permission integrity
package_install
agent_memory
Findings (2)
References agent memory files
SKILL.md ยท code ยท MEMORY.md
Python urllib.request โ network access
resources/check_version.py ยท prose ยท downgraded ยท urllib.request
Why the tier is capped
Execution sink present in raw bytes (Hard Floor: class B). Final tier capped at Caution โ cannot be lifted by any downgrade, example-payload opt-in, or allowlist.
Permissions & capabilities
No declared permissions โ minimal attack surface.
dynamic_evalagent_memorypackage_installcredential_accessnetwork_in Thanks โ recorded.