Guardrails for autonomous tools
The failure mode worth designing against isn't a dramatic one — it's an automated system doing something small and wrong, repeatedly, confidently, with nobody noticing until it's compounded. Heavy-handed guardrails (approval gates on everything, constant confirmation prompts) solve that but make the tool unpleasant enough that people route around it, which defeats the point.
What actually shipped is lighter than that: narrow checks scoped to the specific things that are expensive to get wrong, left alone everywhere else. Alfred's factual answers being kept 100% deterministic — never invented, always traced back to a real note — is the clearest example already live: a guardrail that only exists exactly where a wrong answer would actually matter.
The lesson that generalised past this one project: a guardrail should be able to say what specific failure it's preventing. If the honest answer is ‘just in case’, it's probably adding friction without adding safety.