Caveat verdict
openclaw-smartness-eval
smartness-eval-open-source
Evaluates OpenClaw capabilities across 14 dimensions using local Python scripts; dynamic_eval runs benchmark tests within the evaluation framework.
⚠ Flagged for review — coarse, uncorroborated signal, not a confirmed exploit. Review the config yourself before installing.
Automated static analysis — not a human review. Caveat flags capabilities, not confirmed intent, and can produce false positives. Disagree with this verdict? Use Dispute below.
Findings (7)
subprocess execution — runs system commands from Python
docs/ARCHITECTURE.md · code · subprocess.run(
Recursive delete from root or home — destructive command
config/task-suite.json · prose · downgraded · rm -rf /
Uses eval() — can execute arbitrary code
CONTRIBUTING.md · prose · downgraded · eval(
Pipe to python — executes piped content as Python code
CONTRIBUTING.md · prose · downgraded · | Python
Uses exec() — may execute shell commands
CHANGELOG.md · prose · downgraded · exec(
Python urllib.request — network access
README.md · prose · downgraded · urllib.request
Python os.environ.get — reads environment variable
scripts/eval.py · prose · downgraded · os.environ.get(
Why the tier is capped
Execution sink present in raw bytes (Hard Floor: class B/D). Final tier capped at Caution — cannot be lifted by any downgrade, example-payload opt-in, or allowlist.
Permissions & capabilities
No declared permissions — minimal attack surface.
dynamic_eval Is this flag fair?
Thanks — recorded.