AI

Your AI agent got an A. Did it actually do the homework?

· Anton Ygartua

An orange marble approaches transparent checkpoints on a miniature blue-and-white obstacle course.

A shiny benchmark score is reassuring. It is less useful when your coding agent quietly skips the checks that matter.

In a September 9, 2026 post, Google engineers Taylor Mullen and Christian Gunderman argue for testing the individual actions an agent takes, alongside its overall results. Their examples include checking whether an agent runs a validator after changing a build file and asks for clarification when a request is ambiguous.

Watch the work, not just the score

These smaller checks can help explain why a change improved or broke an agent. The authors describe them as a complement to larger evaluations, not a replacement.

For context, Anthropic’s January guide to agent evaluations also distinguishes an agent’s recorded actions from the actual outcome. A claim that a job is done is not the same as a completed job.

Our take: “finished” should come with evidence. Before trusting a coding assistant with a bigger job, decide what you need to see—tests run, changes inspected, sources checked—and review whether it actually happened.

This is a useful engineering approach, not proof that a particular agent is reliable. Geeknewz has not independently tested the examples.