Research on what AI agent architectures actually cost

Claims about AI agents are usually projections. These papers report measurements: what a given architecture cost per query, how much of that spend reached the answer, and where the numbers stop supporting the argument.

Each one states its method before its result. Query classes, rubrics and thresholds are fixed ahead of a run rather than chosen after seeing the output, and every figure traces to a run record. Where a finding is proposed rather than validated, it says so in the text.

The measurement instrument is specified separately in the AEQ specification, and the system these numbers came from is described in the EAM case study. Written by Michael Valderrama.