dan’s digital workshop
Back to projects

Robin — an agent system that has to earn every node

A small swarm of AI agents, built one node at a time and only when the pain of not having it is real — and the first node turned out to already exist, running quietly since July.

ClaudeObsidianGmailScheduled tasks

Problem

All day I forward links to myself — repos, posts, articles. Forwarding takes two seconds; researching each one properly was the chore, and the reading pile never shrank. The obvious answer is ‘build an AI agent swarm’, and that’s exactly where projects like this die — not from bad code, but from the builder quietly stopping maintenance once seven moving parts outrun a few spare hours a week.

Approach

Before building anything, I ran a premortem on my own project. The top-ranked failure wasn’t technical at all: I stop maintaining it. That reshaped the whole plan. One node at a time, earned by pain — Phase 1 is a single strong Researcher agent, used daily for two weeks before anything else is even considered; boredom isn’t a reason to add complexity. The router comes fourth, not first: my original plan put the orchestrator up front, but an orchestrator with one worker to orchestrate is overhead, so each later node (organiser, verification council, router, content agents) gets a named trigger condition and doesn’t get built until that trigger actually fires. A constitution gives it teeth — local first, $0, every factual claim carries a reachable source or it doesn’t get stated, and irreversible actions (spend, deploy, publish, delete) always need my explicit yes; by voice, only the exact phrase ‘Robin, confirmed’ counts, so a stray ‘yeah’ across the room can never trigger anything. Kill criteria got written down while I wasn’t emotionally invested: a month without wanting to use it, or costing more time than it saves for two months running, or not being able to show it’s better than three months ago — any of those, and it stops, no rationalising around it later. Then the plot twist: before writing the Researcher, I checked what already existed — and found I’d been running one since July. My research inbox pipeline (forward a link to an email address, get a structured, fully-sourced note in my vault by 5pm) was already doing the Phase 1 job every day. So Phase 1 was adopted, not built: a real scheduled task replaced ‘the desktop app happened to be open’, a frozen test set was backfilled from its actual run history, and the one behavioural difference was documented as a named exception in the constitution.

Outcome

Real traffic, daily: 10–14 forwarded threads per run, each verified live rather than summarised from the link — GitHub repos fetched, product claims checked against independent sources, engagement funnels called out instead of clicked. Notes land structured, every claim traceable to a source URL. Items that genuinely can’t be verified — real numbers locked inside an image, say — get filed as needs-input instead of guessed at. The current fight is scheduling reliability, and it’s the platform biting back: the scheduler created triggers its own tools can’t see — invisible to the API listing, ‘page not found’ when opened in the desktop UI. Rather than create a third trigger blind and risk duplicate processing, the observable outputs are now ground truth: did the note land at 5pm, did the Gmail labels move. The bug’s gone to support. Not yet formally scored against the frozen test set — that’s due at the first skill-file change or the fortnight gate, whichever comes first. Then the honest decision: promote node two, or — a perfectly good outcome — add nothing because the Researcher still holds.

Lessons learnt

Design against quitting, not against bugs — the failure most likely to kill this project was never a model or an integration, it was me drifting away, so the ledger, dashboard and kill criteria exist to make drift visible while it’s still cheap to fix. Adopt before you build — the best Phase 1 component had already been running for a month, and naming the exception honestly in the constitution beat rebuilding something worse just to match the plan. And when instruments can’t see the thing, measure the thing — if the tooling can’t show you the scheduled task, watch what the task produces; output is ground truth, instrumentation is a convenience.