What was released
On 29 September 2026, independent benchmarking publisher Artificial Analysis released AA-AgentPerf-Local, a tool for testing the speed of locally hosted AI agents on laptops and workstations. The code, replay workloads, and serving configurations are public. The tool can replay recorded chains of requests against a model server with an OpenAI-compatible interface. For a company choosing hardware for an internal agent, the useful change is that it can repeat the comparison on its own machine rather than rely on a single “tokens per second” figure.
What was actually tested
The initial lineup covers NVIDIA DGX Spark, a workstation with a GeForce RTX 5090, AMD Ryzen AI Halo, and a MacBook Pro M5 Pro. The published tests use four model families in 4-bit versions. The default workload contains eight recorded tasks and 168 model turns, with context growing to about 56,000 tokens. Serving configuration, quantization, and speculative decoding affect results; the authors have published the configurations needed to reproduce them.
According to Artificial Analysis, the RTX 5090 is fastest for the models that fit in its memory. This is not a universal purchasing ranking: a model may not fit in 32 GB of video memory, and purchase price, power use, and concurrent users change the economics. Unified-memory systems can accommodate different model sizes but may respond more slowly in this workload.
The practical boundary
AA-AgentPerf-Local measures latency and throughput, **not answer quality**. By default it excludes the time an agent spends executing tools, and the launch benchmark focuses on a single agent using the entire system. It does not prove that a model will correctly process an order, find a contract clause, or safely update a CRM record. A procurement decision still needs a separate quality test on the company's own tasks, access controls, realistic load, and total cost of ownership.
For a small business, the next step is to choose one repeatable workflow and run the benchmark on available hardware with the model and context length planned for a pilot. Compare task completion as well as speed: if quality is poor, a fast model will not pay for the human corrections it causes.
Photograph of a DGX Spark device included in the benchmark lineup: Daniel Lu / Wikimedia Commons, [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/). For the square cover, the original photo is placed over a blurred continuation of the same image; the derivative is shared under the same license.
