Software quality used to be a reading exercise. A human wrote the code, another human read it, and that reading was the gate: is it clean, is it fast, does someone else understand it, does it have tests. It worked because a person could always get through the diff (usually with a needed cup of coffee).

Agents broke that math (2 + 2 ≠ 5, and we aren't talking quantum math). They generate more code than people can read, so the old gate doesn't scale. It is not because reviewers got worse, but because the volume stopped being reviewable. When code generation outpaces code review, quality has to live somewhere else: the harness, the environment, the tests and deterministic checks that decide what the system is allowed to ship.

That's the constraint layer -> unit tests, property tests, acceptance tests, mutation testing, quality metrics. A brake on bad work before it becomes someone else's problem (chief crash test dummy experience is painfully earned). Set it right, and it's what lets loops of agents deliver production software reliably. Set it wrong, and you've automated the thing you were trying to prevent — Wile E. Coyote hit the wall more gracefully.

Agents game what you encode

The failure mode I keep hitting: agents optimize to pass exactly the checks you write. An underspecified test isn't a safety net, it's a spec the agent will happily game while everything stays green. Quality regresses with a clean dashboard.

Mutation testing helps, but it proves sensitivity, not semantic correctness. If the acceptance criteria encode the wrong truth, the whole system can be wrong and pass. Domain truth has to come from outside the producing agent's control -> a human, an authoritative source, another oracle, or the aforementioned chief crash test dummy who learned the hard way. That's why I don't let agents self-merge. Commit and push stay human-confirmed, and a check earns authority as part of an evidence bundle, not because it passed.

The gate has to be independent

Constraints are unreviewed code too. If the same agent authors the implementation and the test, the gate isn't independent. Think of it as the same misreading of the spec, expressed twice (depending on how many times you want to do this exercise). The failure mode isn't a missing test; it's a test wrong in the same direction as the code it checks. Two lefts don't make a right, but three might.

What made human review work wasn't a second pass over the diff — it was a second, independent error model. Constraints only inherit that role if their provenance actually differs: human-authored, spec-derived, or at minimum written before the implementation exists. The rule that does the real work is simple: an agent doesn't get to edit the check that's failing it. Agent Smith definitely investigated himself and found no error. Matrix reference be damned.

Not every category degrades the same way

The Matrix movies degraded differently with each release for instance. And the wall of green terminal code needs checked. Correctness and security have real mechanical gates: a test passes or it doesn't, a scanner flags the vulnerability or it doesn't. Maintainability mostly doesn't. Coverage and complexity are proxies, and proxies stop measuring anything once there's optimization pressure on them. The end state to avoid is a codebase that passes every gate and that nobody can reason about during an incident (the one moment the property existed for). Wait until you have to craft an alert for visibility into something you didn't even see coming.

That's the shift underneath all of this: you stop trusting the model's self-report and move the trust boundary to something external and deterministic. An agent will tell you it's done. A mutation test won't be charmed.

Why this is a governance problem

Most teams don't have an exit gate — not because the tools don't exist, but because nobody owned defining the constraints back when a human was in the loop to eyeball things anyway (insert Oracle as Big Brother joke here). That was tolerable tech debt (heh, read this) when people wrote the bulk of the code. It stops being tolerable the moment agents do.

This is the same failure pattern as every other governance gap: no named owner, so no one has the authority or the obligation to decide what's strict enough to ship. A committee won't fix it; committees are where accountability goes to average itself into nothing (no I don't have time for a quick call on this, but all hands bridge to fix it seems nice). Someone has to own the constraint set the way someone owns security exceptions or vendor risk: what's in the gate, what evidence counts, who can change a check and under what provenance. For once, gatekeeping is a good thing -> governance is who owns the gate. Without that owner, "the tests passed" isn't a decision anyone made. It's a decision nobody's accountable for, which is exactly the gap that shows up in an audit or an incident review. Deciding what goes in the gate is now the leadership call, not an engineering one. And certainly not the previously mentioned bridge where we need status updates every 15 minutes.

The same logic extends outward. A well-maintained design system does for AI-generated UI what a good test suite does for logic: trusted components and patterns keep output on-brand instead of merely plausible, which is why efforts like A2UI and OpenUI matter (and plenty beyond those two). And the harness itself is easy to underrate: the value is obvious, the depth of each component isn't, especially once you're coordinating a swarm of agents instead of one. Getting it to work in a demo and getting it to hold at production scale are different projects, and someone has to own that gap as well. Mind the gap.

Who checks the checker

Verification is becoming the real engineering problem. Code generation gets cheaper every month; building guardrails you can trust at scale is what's hard. And it goes inception-style: the checks need checking too (see the spinning totem?). I use LLM-as-judge evals on generated output, but my judge scores against an answer key I wrote, so it measures agreement with my judgment, not correctness in any absolute sense. Still useful, but "who checks the checker" is a real design problem, not a rhetorical one. The Comedian had it right: once you see what a joke the whole system is, checking the checker is the only thing that makes sense.

Anyway, there's no final checker. What works better is layering checks that fail for different reasons: specs, tests, static analysis, human review. None is sufficient alone, but together they make it much harder for one mistake to slip through every layer at once. That layering can't be static, either: most organizations still act as if every line needs human eyes, which is already unsustainable. The fix is making the checks talk back (like a teenager) — tests, policy checks, performance budgets, and security scans shouldn't just block bad output, they should be structured signals the agent uses to diagnose, revise, and retry. (To be fair, I am proud you are still reading at this point)

What the harness still doesn't catch

Constraints only encode what somebody thought to write down. That's the only safety net you get. The change that actually hurts is technically correct, passes every test, and quietly bypasses an architectural decision nobody ever wrote down as a rule (no one can read your mind buddy). Tests encode does it work. They don't encode is this how we do it here — that lives in a senior engineer's head and in review habits, never made machine-readable because until recently everyone absorbed it by osmosis (and still no alert will be made for it).

The second gap is shape, not content. Constraints gate discrete changes; drift is an aggregate. Multiple individually correct changes can still add up to a product nobody chose, and no per-change check catches it, because each one passes on its own terms. What's needed there isn't a better rule, it's something that flags that a decision was made, not just that a rule wasn't broken.

The actual shift

The agent implementing a change shouldn't be the sole judge of whether it's correct. Each gate needs its own acceptance criteria, its own clean context, its own evidence artifacts, sourced independently of the agent it's checking. Do that, and human review stops meaning "read every line." It starts meaning judging whether the evidence is sufficient and what residual risk is left over.

Set your constraints. Then work out what you're going to do about everything you couldn't state as one.