15/08/2026
Everyone can build an agent now. Almost nobody can tell you whether theirs is getting better.
You reword the prompt, swap the model, add a tool. It answers — fluent, confident, plausible. Did that change help?
You can't say. Last week's incident is gone: the deploy is old, the pool is back to normal, the logs rotated. So you skim two outputs, decide it "seems better," and ship. That's tuning by guesswork, and most of us are doing it in private.
The fix isn't a smarter model. It's the test suite tool-using agents never had:
1. Record one case whose answer you already know.
2. Replay it to the agent — same data shapes, frozen scene.
3. Score what it did, not how it sounds.
Four checks do most of the work: right cause, right evidence, rejected the planted distraction, stayed inside the step budget.
The evidence one matters more than it looks. An agent that names the right cause but never opened the deploy log got lucky — and luck doesn't survive the next incident. Grading only the final answer can't tell those two apart.
I'm writing the whole thing up as a short book: a small evaluation harness in plain Python, no framework, built end to end on a DevOps incident agent and ready to point at your own.