Writing
Field notes on agentic engineering — harnesses, orchestration, and what it's like to stop reading your own code.
The Finish Line Keeps Moving
Before TutorPro's soft launch, I had agents rebuild its help center from the source code. They found a tutor who couldn't accept a parent's request while 13,758 tests passed, and every pass since has found more. Quality is still the hard part of agentic development, and I want your ideas for closing the gap.
Feeding the Factory: One Paragraph In, Three Tasks Queued
A trace through one real session: the paragraph I typed, the standing rules and memory already loaded before it, the lint that gates every spec before an issue exists, and the three dependency-ordered agent tasks that came out the far end. Both skill files are linked as public gists.
I Built My Own Claude Tag
Anthropic put Claude in Slack. I already had Claude running on a Mac mini with my tools, my checkouts, and my task queue, so I gave that one a Slack handle instead. Now I @mention composer with anything, it remembers the thread, and this week it talked me through my own stale rule before pulling a repo.
Time, Not Safety
Two of the most powerful people in computing published essays weeks apart asking the world to slow down or brace for impact. I read them as someone who was terrified of losing his job to AI, bet on himself, and landed on his feet. The bet bought time, not safety.
The Cheap Model Was Fine. Proving It Was the Hard Part.
I built a gateway that sends each prompt to the cheapest Gemini model that can answer it, then benchmarked it on 300 prompts with a Pro-tier judge. The cheap model held up. My test set, my verifier, and my first table of numbers did not. Six lessons for anyone paying an AI bill.
Say It Out Loud: A Talking Tutor and the Interface That Disappears
I gave my AI Labs Study Buddy a voice. Talking Tutor quizzes you out loud on your own notes over the Gemini Live API. Building it changed how I think about where voice fits, and why the next wave of AI interfaces might not have a screen at all.
The Catalogue Tax: Why I Still Hand My Agents a CLI
Every MCP server you attach re-sends its tool catalogue on every request, and a new controlled study shows most agent harnesses pay that tax eagerly. Here's what MCP really costs in tokens and speed — and why I still give my agents a CLI when a good one exists.
Dangerously Convenient: My Agents Learn to Order Breakfast
DoorDash shipped a CLI built for AI agents, so I pointed two of them at it — Claude Code and Antigravity — and asked for breakfast. What happened next says a lot about how agents are going to consume the internet.
Teaching the Factory to Pick Its Own Tools
My AI dev platform was choosing models with a regex. A $54 accounting bug, one research paper, and a morning of planning later, Composer is building its own complexity-aware model router. Here's the design.
In Memory of Lucy
A poem for my beloved dog Lucy
The Inverse Architecture: Building the System That Builds the Systems
Four months ago I introduced Composer, my AI software factory. Since then it's shipped 642 pull requests for about $4K — and I've stopped reading the code. Here's the architecture everyone else is chasing backwards.
Hold It Loosely: What AI Is Teaching Me About Letting Go
After 30 years in software, I'd built my identity on the things I made. AI is teaching me to hold them loosely — to build, ship, and let go — because the half-life of cleverness has never been shorter.