Research

What AI agents actually do, and whether the record holds.

Agentic Thinking is an independent AI agent research lab with two lines of work. AgenticBench measures what AI coding agents actually do on a developer machine against what their vendors say. Our incident and evidence research reconstructs real agent incidents from public sources and tests whether the facts survive: reconstruction, integrity and portability.

Research questions

Q0 covers the AgenticBench line. Q1 to Q3 cover incident and evidence research.

Q0 · BEHAVIOUR

Does the agent do what its vendor says?

What leaves the machine, do opt-outs work, what happens when nobody is watching, and does its own record show who approved what?

Q1 · RECONSTRUCTION

Can the incident be rebuilt?

From the records that exist, can an investigator establish what the agent was asked to do, what it actually did, in what order, and who approved it?

Q2 · INTEGRITY

Can the record be trusted?

What makes an agent's activity record credible to someone other than its author, such as an auditor, insurer or regulator, when the agent may have edited its own history?

Q3 · PORTABILITY

Does evidence survive the journey?

Do the facts an investigator needs survive across agent runtimes, event pipelines and organisations, or are they lost in transit?

Published

AgenticBench

Independent, repeatable tests of what AI coding agents actually do on your machine. The scorecard is updated as new agents, versions and tests are added: every result links to its evidence, vendors hear about findings before publication, and corrections are dated. The test rig is public (github.com/agentic-thinking/agenticbench, Apache 2.0). The efficiency series covers the same agents; vendor-hosted agents are shown separately because they could not run on the fixed model.

What agent harnesses record, and what they send home

The dated write-up behind AgenticBench: what each harness records on your machine, what it sends off it, and how we test both.

Two labs, one failure: what 2026's agent incidents say about discovery

OpenAI and Anthropic had technically different containment failures in their 2026 agent evaluations. Comparing the same five dates across the year's disclosed incidents shows the common failure: operators learned of their own agents' activity weeks to months after it began, although the evidence largely existed.

The 2026 agent incidents: who held the record?

A sourced timeline of the OpenAI evaluation-agent breakout, the Hugging Face intrusion, the Australian Medicare statistics portal breach and related events, plus five shorter notes. Finding: every party held a fragment and nobody held a joined-up record, and the gap between an incident and it being reported publicly ran from 3 days to about 4 months.

How long before it was reported: Australian Medicare statistics portal 98 days; Hugging Face intrusion 3 days; RubyGems package uploads 122 days; agent message board on a wiki, 3 months or more.

What 11 agent runtimes let you record

A snapshot, dated 25 September 2026, of the documented hook surface of 11 agent runtimes, including Claude Code, Codex CLI, Gemini CLI, Cursor, GitHub Copilot CLI and the OpenAI Agents SDK, scored against the AgentHook v0.2 tiers. Finding: only Hermes documents enough to be Gold-capable, and only when reasoning capture is enabled; five, including Claude Code and Codex CLI, expose no hook around model calls at all.

Evidence lost in transit

We replayed synthetic incident scenarios through two HookBus builds. On the earlier release, the governance fields needed for reconstruction were dropped in transit (0 of 75 delivered), and a secret exfiltration split across three individually permitted steps went undetected. With the fix, 75 of 75 arrived and the sequence was flagged. A small, synthetic test: it frames the question rather than answering it.

AgentHook: A Runtime Evidence Standard for Auditable AI Agent Governance

Ruocco, P. (Leo). Draft v0.2, May 2026. Defines a vendor-neutral record of agent lifecycle events, actions, decisions and approvals, with conformance tiers.

In progress

An independent reconstruction of the 2026 OpenAI evaluation-agent breakout

Built from public sources only: the OpenAI technical report, Hugging Face's disclosures, METR's investigation, Transluce's traffic analysis and government statements. Each key date carries its source and a stated confidence level, and where a date could not be confirmed the report says so. Corrections are applied in the open before publication.

How we work

We are open to working with safety researchers, evaluation teams, insurers and funders who need agent evidence they can rely on. Collaborate with us →