Agentic Evals

Our evaluations now live at Agentic Evals.

Agentic Evals is the evaluation service of Agentic Thinking. If you build an AI agent runtime, or a product where an agent takes actions through tools, APIs or shells, we run it, try to break its guardrails, and tell you plainly what holds and what doesn't, with evidence. Every evaluation is led by a person, not a script.

Separate from AgenticBench

AgenticBench is separate: it never accepts funding from the agents it tests, and coding agents on the benchmark are not eligible for paid evaluations.

Declared interests

Declared interest: we maintain the open-source HookBus project and the AgentHook standard. We sell no agent-governance product.

We have committed to transferring the AgentHook standard to an independent foundation. Our full list of declared interests is on the Trust page.