๐Ÿงฎ Part 3: In Bots We Trust

A teammate told me he trusts a pull request more now that it's passed his review, plus a couple of agents built to catch what he might miss. Three layers of review sounds like safety. It might just be a receipt for a problem nobody fixed.

๐Ÿงฎ Part 3: In Bots We Trust

"It depends."

That's what I said in the retro. My manager had asked, point blank, in front of the team: how helpful has Mauve been? I'd already spent weeks writing up detailed notes and answering surveys โ€” what broke, what worked, what needed fixing โ€” about the new autonomous agentic development tool and at that point I didn't have the energy to say it all again out loud. So I gave the short version and left it there.

A teammate pulled me aside right after. No really, what do you actually think?

Here's the actual answer, the one that didn't fit in a two-word retro response: it depends on what you're using "confidence" to mean. Because we're placing more confidence in artificial intelligence than actual intelligence. Because we've placed more confidence in the output of really intelligent pattern recognition. Because we've started measuring confidence in a pull request by how many separate reviewers, especially non-human, signed off on it โ€” and I don't think that number means what we think it means.

Quick context if you're not in the weeds day-to-day: a pull request is just a request to have your work checked before it becomes part of the real thing โ€” code, in this case, but the concept is the same as sending a draft to an editor before it publishes. Nothing goes live until someone else looks at it and says yeah, this is good.


๐Ÿง‘โ€๐Ÿคโ€๐Ÿง‘ Redundancy Isn't Confidence

These days, that "someone" is rarely just one someone. It's a teammate's eyes on the change, plus one or more agents โ€” bots built specifically to review code โ€” layered on top, each looking for something different. My teammate's version of confidence was: now that it's gone through my teammate, and the review agents, I feel better about it. And he wasn't wrong that it feels safer. Feeling safer and being safer aren't the same thing, though, and the gap between them is the whole point of this post. Feeling safer can be dangerous.

Here's the reframe: when a system needs three independent reviewers to trust its own output, that's not redundancy working as designed. That's redundancy working as a symptom.

A teammate's review catches judgment calls โ€” the stuff that requires actually knowing the codebase and the person who wrote the ticket. One review agent might catch spec drift, assuming nobody already approved the drift, in which case it catches nothing. Another might catch a known category of bug pattern. Each layer is looking for something different, which sounds thorough โ€” until you notice that none of them is actually validating the whole. They're stacked, not overlapping.

Think of it like running water through three separate filters. Each filter is designed to catch something different โ€” sediment, then chemicals, then bacteria. Run it through all three and the water tastes clean. But if you genuinely need three filters before you're willing to drink it, the real question isn't "did it pass?" It's what's wrong upstream that you don't trust one filter to catch on its own?


๐Ÿ’ธ Diminishing Returns, Now With Robots

This is the part that should sound familiar, because it's not a new problem โ€” it's an old problem wearing an AI costume.

Throwing more review agents at a pipeline is the same move as throwing more money or more people at a problem. It works, for a while. Then you hit the wall every scaling curve hits eventually: marginal reviewer, marginal return. The third layer of review doesn't catch three times what the first layer catches. It catches slightly more, at meaningfully higher latency, for a PR that was already slow to land.

In other words: the problem was never a shortage of reviewers. It was never a shortage of intelligence โ€” artificial or otherwise โ€” looking at the diff. Bolting another layer onto a review pipeline that's already too slow solves for a problem you don't have, while quietly ignoring the one you do.


๐Ÿ”ง What Actually Buys Confidence

If three layers of review are genuinely required to trust a change, that's not evidence the system is working. That's evidence the system has a problem three layers of review are compensating for โ€” and compensating for a problem isn't the same as fixing it.

A team that needs a teammate and multiple agents to feel good about a merge has a much shorter list of actual fixes sitting right in front of it:

  • โœ… Tests that actually cover the thing that keeps breaking, instead of tests that exist to satisfy a coverage number
  • ๐Ÿ—๏ธ Architecture clean enough that a change in one place doesn't require three people to reason about the blast radius
  • ๐Ÿ“ Following the frameworks and conventions the team already agreed on, instead of every service reinventing its own patterns
  • โœ‚๏ธ A smaller codebase โ€” fewer features nobody uses, fewer dependencies nobody remembers adding
  • ๐Ÿค Less siloing, more building as an actual team, so "confidence" isn't something you have to manufacture after the fact because nobody but the author understood the change going in

None of that is exciting. None of it involves a chatbot. All of it actually moves the needle on whether a change is trustworthy, instead of just making it feel trustworthy because more eyes โ€” human or otherwise โ€” glanced at it before it merged.

You don't need a third reviewer. You need to ask why the first two weren't enough.


Final Thought ๐Ÿ’ญ

So is more review better? Sure, if you're optimizing for how safe a merge feels in the moment.

The harder question is whether that feeling is buying you anything real, or whether it's just latency with better packaging. Three layers of review didn't make the code more correct. It made the team more comfortable shipping code none of them fully trusted on their own โ€” and comfort, at scale, is expensive.

Confidence you have to buy in triplicate isn't confidence. It's a receipt for something nobody actually fixed.

Every day, my confidence in the few making decisions impacting the many is waning. I don't need redundancy to be confident about that.