Small agents beat big ones, most of the time
The instinct when building with AI agents is to reach for one capable, general agent and give it everything — it feels more efficient than managing several narrow ones. In practice, the general agent is the one that's hardest to trust, precisely because it's hard to say in advance what it will and won't do in any given situation.
A narrow agent with one clear job is easier to reason about from the outside. If it only does one thing, checking whether it did that thing correctly is a small, bounded question. If it does twenty things, checking its output means checking twenty different kinds of correctness at once, and most of the time that check just doesn't happen — the output looks plausible, so it gets trusted, and plausible isn't the same as right.
The trade-off is real, not free: more small agents means more coordination, more places for a handoff to go wrong, more surface area overall. The failures just move from ‘wrong inside one big system’ to ‘a coordination gap between two smaller ones’ — which is a better place for them to live, because it's visible, but it isn't nothing.
The dividing line that's worked so far: split by what needs to be individually verifiable, not by what feels like a natural feature boundary. A step that produces a factual claim, touches something irreversible, or is genuinely hard to check by eye is a strong candidate for its own narrow agent. A step that's just formatting or shuffling data usually isn't worth the coordination overhead of splitting out.
Small agents beat big ones most of the time — not always, and not because small is a virtue in itself, but because narrow scope is what makes trust actually checkable instead of assumed.