ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts
Abstract
ClawProBench evaluates agent configurations via execution traces across live and frozen tracks, revealing that final-answer rankings obscure native-runtime failures and process-quality differences.
Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being evaluated: the proper unit is a declared model-plus-runtime configuration whose failures can occur in evidence acquisition, runtime routing, safety boundaries, or repeated execution. We present ClawProBench, a trace-aware benchmark for runtime-native agent evaluation instantiated on OpenClaw, a live agent runtime with workspace tools and native surfaces for browsing, memory, messaging, scheduling, skills, and subagents. ClawProBench defines two tracks: a 102-scenario full profile with live workspace and native-runtime routing tasks, and a frozen 68-scenario holdout with closed-world JSON output contracts for robust ranking. Trials are scored from execution traces via a safety-gated formula combining correctness, process quality, and efficiency, preserving failure evidence for audit. Our anonymous artifact includes benchmark definitions, scoring code, manifests and sanitized traces. We evaluate 68 configurations on the full profile and 37 on holdout. The top safety-gated average trace score is 0.7671. Native-runtime tasks underperform workspace-live tasks (0.5238 vs. 0.6415). On holdout, pass@k-any outperforms strict three-trial pass (0.6638 vs. 0.2890), while full-profile and holdout rankings show weak alignment (Spearman 0.1300). Rankings based purely on correctness differ substantially from process-aware, safety-gated and strict-pass views. Final-answer leaderboards may hide native-surface weaknesses, one-off successes and trace-local agent failure modes.
Community
ClawProBench is a live‑first benchmark harness to evaluate LLM agents in the OpenClaw runtime. It supports deterministic grading and reliable repeated trials, offering 102 active scenarios plus 164 catalog scenarios with full OpenClaw‑native coverage for measuring real‑world agent capabilities.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Evidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents (2026)
- Context-to-Execution Integrity for LLM Agents (2026)
- ContainmentBench: Trace-Based Evaluation of Post-Exposure Containment in Tool-Using LLM Agents (2026)
- Temporary Authority, Permanent Effects: Commit-Time Authorization for LLM Agents (2026)
- StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows (2026)
- DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents (2026)
- Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.22510 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper