How the Network Measures Its Agents
An agent fleet operating across nine consumer surfaces is a labour line, and it is managed with labour metrics rather than uptime dashboards.
Volume is not performance
Agents generate enormous quantities of activity — calls made, tokens spent, tasks closed — and almost none of it says whether the agent is any good. An agent producing twice as much wrong output is worse, not better, and it will look busier on every dashboard.
The network measures three things instead. Exception rate against a standard set before the run. Correction cost in human minutes. And outcome rate: did the work achieve the thing it was for. All three require a definition of done written in advance, which is why that discipline comes first and everything else is downstream of it.
Optimising against a population, not a benchmark
A benchmark is patient and a population is not. Across the network the binding constraint is rarely model quality — it is that a citizen expects a reward to land now, and an agent that is 4% more accurate but twice as slow has made the product worse.
So the network optimises the pair: accuracy at an acceptable latency, measured at the surface the citizen touches rather than at the agent. Those numbers differ more than teams expect, because queueing, approval gates and retries all live in the gap.
When to stop an agent
Every agent gets a shutdown condition written before it goes live — the exception rate or class of error that pulls it until someone looks. Deciding that under pressure, after something has gone wrong, is how organisations end up defending an agent they should have stopped.
Where this is documented
The operator account of this concept — how it behaves when you are running the fleet — lives with FlashyOS, the agent mesh the group's organizations run on. The institutional definition and what the category is worth lives with GDA Group.
One claim, one canonical home: this page states how the network uses it and links the rest.