Caveat verdict
perf-test-flagos
vLLM performance benchmarking inside containers against a user-controlled model; all network interaction is internal (localhost:8000) and behavior matches the stated benchmarking purpose.
โ Flagged for review โ coarse, uncorroborated signal, not a confirmed exploit. Review the config yourself before installing.
Automated static analysis โ not a human review. Caveat flags capabilities, not confirmed intent, and can produce false positives. Disagree with this verdict? Use Dispute below.
Permission integrity
network_out
Findings (3)
Pipe-to-python pattern โ remote code execution risk
SKILL.md ยท code ยท curl -s http://localhost:8000/v1/models | python
Pipe to python โ executes piped content as Python code
SKILL.md ยท code ยท | python3
subprocess execution โ runs system commands from Python
scripts/run_all_benchmarks.py ยท prose ยท downgraded ยท subprocess.run(
Why the tier is capped
Execution sink present in raw bytes (Hard Floor: class C/D). Final tier capped at Caution โ cannot be lifted by any downgrade, example-payload opt-in, or allowlist.
Permissions & capabilities
No declared permissions โ minimal attack surface.
network_outnetwork_in Is this flag fair?
Thanks โ recorded.