Meta / Contents
Feeding the Factory: One Paragraph In, Three Tasks Queued
A trace through one real session: the paragraph I typed, the standing rules and memory already loaded before it, the lint that gates every spec before an issue exists, and the three dependency-ordered agent tasks that came out the far end. Both skill files are linked as public gists.
Here is everything I typed this morning:
/compose-workI'd like to add @ mention tagging in group/contest/leaderboard message board. This should only allow user to tag users in the contest/group depending on context. Also, when users click reply, the message should start with an @ mention of the user they are responding to. When a user is tagged, they should get an app notificiation and push notification - the push notification should navigate directly to where they are mentioned, so we may need to also add anchors in web and something in mobile to scroll to that part of the board.
One paragraph. One typo. No file paths, no API contracts, no mention of which of the three apps in that monorepo would need to change.
What came back was three cross-linked GitHub issues, numbered 1897, 1898 and 1899, plus three agent tasks sitting in a queue, with the two client phases waiting on the backend phase.
Back in December I wrote about the spec-writer that interviews me before I'm allowed to build anything. That post was about the conversation. People have been asking me the next question since: okay, but what's actually under there?
This is that. One real session, every layer, both skill files linked at the bottom so you can take them.
What was loaded before I typed a character
The paragraph was short because everything else had already been written down.
Three things were in context before the Claude Code session opened, and all three changed the output.
My global instructions. ~/.claude/CLAUDE.md is a symlink into my skills directory. That's deliberate. The skills directory is a git repo, so my standing rules get versioned alongside the skills that depend on them. Rules like feature branches for all code changes, checkpoint before committing to a path, and always use spec-writer.
The project's instructions. A CLAUDE.md at the repo root that points at AGENTS.md, which points at a per-app AGENTS.md under each of the three apps in the monorepo. The root file also carries a standing rule that every admin feature has to extend the admin tooling in the same PR. I wrote that one after shipping three features that quietly didn't.
My memory index. Roughly seventy one-line pointers, each linking to a file holding exactly one fact. This is the part that does the most work and gets the least credit, so it gets its own section later.
None of that is typing I did this morning. It's typing I did over months, and it's sitting there every time.
The skill chain
Three of my own skills ran, nested:
/compose-work the orchestrator
└─ /spec-writer the quality bar
└─ /grill-me (available, not invoked this time)
compose-work is the outer loop, and its opening principle is the thing that makes the whole setup livable:
Key principle: Front-load ALL user interaction, then execute autonomously. Gather every decision (requirements, task type, model, effort) before creating issues or tasks. Once the user gives final approval, the entire creation pipeline runs hands-off.
That rule exists so I can leave. Answer the questions, approve the plan, walk away. Issues get created and tasks get queued without me. I have a hook that plays a sound when the session needs me again, which sounds like a gag and is genuinely load-bearing. Front-loading only pays off if something tells you the front-loading is over.
spec-writer is where the quality bar lives, and I'll get to it in a second.
One honest note: grill-me is wired into spec-writer as the interview tool, but it didn't run as a separate skill here. The interview happened inline instead, through the multiple-choice question cards. Same effect, one less hop.
Eleven decisions, three rounds
The interview was eleven decisions across three rounds. Every one arrived as a multiple-choice card with a recommended option and the tradeoff stated out loud.
That framing matters more than it sounds. I wasn't asked to describe what I wanted. I picked from options that had already been checked against the code.
Round one covered how mentions get stored, who's mentionable, what the reply prefix actually means, and notification dedup. Round two covered the @everyone permission gate, what happens when two people share a public display name, where a deep link lands, and what an edit that adds a mention should do. Round three was phase structure, task type, and model plus effort.
Two of those are worth showing, because they're the ones I'd have gotten wrong alone.
The storage format question didn't arrive as "how should we store mentions?" It arrived with the fact that installed mobile builds render message content verbatim, which is what made plain text plus a parallel array the recommended option instead of an inline markup token. Ship the token and every phone that hasn't updated shows readers the raw markup.
The dedup question came with a line number. The message service already notifies the parent author when someone replies. Add a reply prefix that mentions that same person, and they get pinged twice for one reply. The question wasn't do you want dedup, it was here's the collision, pick the suppression rule.
I wasn't thinking about double notifications when I wrote that paragraph. I was thinking about the feature.
The gap between the feature I imagined and the system it lands in is the entire value of the interview.
Specs aren't written for me anymore
This is the biggest change since the December post, and it's one line near the top of spec-writer:
Audience: the specs are consumed by LLM implementing agents (of varying models and capabilities), not humans. Different models fill spec gaps differently — wherever a spec is ambiguous, each implementer invents a different answer. Precision converges them; narrative prose does not.
Sit with the middle clause. Each implementer invents a different answer. A human developer who hits a vague spec asks you about it. An agent fills the gap confidently and keeps going. You find out in code review, or you don't.
Every rule in that file follows from that one sentence. Four of them do the heavy lifting:
Interface-first precision. Any new or changed endpoint, DTO, config property, schema, collection, CLI, or event is stated EXACTLY: path, name, type, default, error codes/status codes. Ambiguity in contracts makes implementations diverge; ambiguity in internals is fine — grant it explicitly: "Internal design is implementer's choice provided the stated contract holds."
Context split rule. Include everything non-discoverable from the repo: decisions made during the interview, policies, cross-system gotchas, production knowledge. Point to (file path/symbol) rather than duplicate what's greppable — implementing agents research the codebase themselves.
Verification-first. Every acceptance criterion must be objectively checkable: an exact command, a named test that must exist and pass, a request→expected-response pair, or a greppable condition.
Traceability. Every requirement maps to ≥1 acceptance criterion; no criterion exists without a requirement behind it.
Principle 1 is the one people skip. Being vague about internals is fine. Better than fine. It's respectful of an implementer that may be smarter than you about the internals. Being vague about the wire is how you get three phases that don't fit together.
Principle 2 is the one that saves the most tokens. Don't paste the file; name it. The agent can read.
The lint that runs before anything is created
Here's the part I'd actually hand to someone copying this.
spec-writer has a step labeled MANDATORY GATE. Before a single issue is created, every draft spec gets scanned:
### Step 5: Ambiguity Lint (MANDATORY GATE)
Before any issue or file is created, lint every spec draft. A draft failing any
check goes back to Step 3/4 — never ship the ambiguity to the implementer.
1. Banned undefined qualifiers: "appropriate", "properly", "correctly",
"gracefully", "robust", "as needed", "if necessary", "handle errors"
(without defined behavior), "etc." / trailing open enumerations.
2. Pattern citations: every "follow existing pattern" names a file and symbol.
No uncited "as done elsewhere".
3. Complete enumerations: lists of endpoints, fields, screens, or cases are
complete, or the rule generating them is stated precisely.
4. Contract exactness: every wire-visible or persisted artifact appears with
its exact name/path/type/default/error behavior.
5. Traceability: every requirement has ≥1 acceptance criterion; every
criterion traces to a requirement.
6. Checkability: every acceptance criterion is machine-checkable or a named
observation. "Works correctly" never survives the lint.
7. Non-discoverable context present: interview decisions, policies, and
gotchas that cannot be found in the repo are in the spec.
8. No running-environment verification for agent-executed specs.
Rule 1 is a grep. Rule 8 is a grep. That's the whole implementation. Two grep -nE calls across the draft files, run before the GitHub CLI fires gh issue create.
That shape is on purpose, and it's the same one I described from the inside in The Inverse Architecture: the judgment belongs to the model, the gate is boring deterministic code, and nothing gets created until the boring code says so. A model that writes a spec is also a model that can talk itself into a vague one. A grep can't be talked into anything.
It caught exactly one violation this session: the word "correctly" in a manual-check step on Phase 3. The fix wasn't deleting the word. It was turning the sentence into something a person could actually fail:
Before: verify the board scrolls correctly to the mention.
After: on a thread of 50+ messages with the mention in the first 5, confirm the board lands centred on it.
Same intent. One of them is checkable.
Rule 8 deserves its own note, because it's the rule most specific to my setup and the one most likely to matter in yours. Specs that get handed to an autonomous agent may only verify with the test runner, the build, and greppable conditions. Never npm run dev, never a curl against a port, never a database outside a temp directory the test made itself. Anything needing a live app goes in a separate section addressed to a human on a local machine. The agent is told, in its own prompt, to skip that section.
If you hand work to agents running on a host you don't fully control, write that rule before you need it.
The time it caught itself
spec-writer has a step requiring me to be asked about the live UI before any UI phase gets spec'd. This session skipped it. The templates and components got read directly instead, and the question was never put to me.
My first explanation for why was wrong, which is the interesting part. I assumed the no-live-server rule had made the running UI off-limits. It hadn't. That rule governs what the implementing agent may do on a host I don't control. It says nothing about what the spec author may do on my own Mac.
So the step was just skipped, on a feature that adds a picker, a highlight flash, and a reply prefill across two different clients. It's now a hard gate with a closed list of valid reasons to skip it, and this line at the bottom:
"I read the templates/components instead" is NOT a valid skip reason. Static reading gives you structure and class names. It does not show layout pressure, cramped containers, what a real empty or loading state looks like, or whether a new control has anywhere to go. Those are exactly what a UI phase's spec has to get right, and they are the things a spec author invents plausibly and wrongly.
Reading a 1,200-line component tells you its structure. It does not tell you whether an autocomplete menu has anywhere to go on a small screen.
A system you never audit doesn't improve. It just accumulates.
What memory actually changed
This is the part that surprised me when I traced it back. Eleven memory entries measurably altered these specs:
| The remembered fact | What it changed |
|---|---|
| Canonical links only render on builds ≥ 1.3.0 | Made plain text plus a parallel array the recommended storage format |
| The public name display rule | Fixed the mention display name, and drove the whole name-collision question |
Never use a plain task type for code or PR work |
Removed an option from the task-type card entirely |
| Overlapping tasks chain; independent ones don't | Made the two client phases wait on the backend instead of all three firing at once |
| A same-field derived query throws at execution time | Went into Phase 1 as a named gotcha, with the fix |
| An ID-field mismatch is unproven and needs real containers | Went into Phase 1 as a warning against verifying it with a mock |
| Web is on FontAwesome 5.15.4 | Phase 2 gotcha. Icon names added in FA6 render as nothing |
| A banned spacing class, and a dark-mode colour collision | Phase 2 gotcha, with the carve-out explained |
| A known double-mount bug on nested routes | Phase 3 gotcha. Don't push a route that's already mounted |
| Bottom-sheet scrolling needs two specific style rules | Phase 3 constraint against restructuring the modal to fit the autocomplete |
| No releases Thursday through Sunday | Why the app-store release got flagged as a separate Monday-to-Wednesday step |
Read that column on the right again. None of those are things an agent finds by grepping the repo. They're the residue of previous sessions, one fact per file, and they're why these specs read like they were written by someone who'd already shipped in this codebase.
The loop closes at the end, too. A new memory file went in recording the three issue numbers, the three task IDs, and every product decision from the interview, with a note to check the issues are still open before acting on it. One line went into the index pointing at it.
Every session pays a little rent forward.
Queued, then hands off
The last step queued three agent tasks, one per issue, and made exactly one interesting decision.
Both client phases wait on the backend phase. They're siblings, though, not a sequence. Web has no reason to wait on mobile. So the tasks were created individually with a dependency on the backend task, rather than as a chain, which would have serialized web behind mobile for nothing. My standing rule only calls for a chain when tasks genuinely overlap:
Add
depends_ononly when a phase consumes an artifact or code contract delivered by an earlier phase (the issue says which and why), or when the user asked for a chain.
One more small thing that saves me constantly: the model options I get offered are never hardcoded. compose-work queries the current model list at run time, because hardcoded names go stale. Live quota comes back with it, which is why the recommendation I saw that morning mentioned my seven-day window sitting at 57%.
From there it's out of my hands. Composer picks the tasks up, builds, tests, reviews, opens PRs, and deploys to dev. I've written about that side of it twice already, first with OpenClaw and then with Composer itself. This post is about the other end of the pipe.
The chain stops at queued deliberately. Implementation, review, merge, and deploy are separate skills with separate gates, and two of those gates are standing rules no agent can clear alone: never self-merge an unvalidated PR, and never release during a live tournament.
I'm not orchestrating the build
People ask me what to call this, and I've never had a clean answer. Business analyst? Product manager? Neither fits.
In December I signed off that post saying I felt like "a senior architect orchestrating agents." Nine months on, I'd drop the second half.
I don't orchestrate the construction. I don't watch the agents work, I don't sequence their steps, I don't review their intermediate output. I find ideas, I describe them precisely to a setup that knows my codebase better than my memory does, and work becomes the output.
That's just being the architect. You draw the thing, you specify it to a standard that can't be misread, you sign it, and then a crew you're not standing over builds it. The drawings have to be good because you won't be there.
Which is why the skills at the bottom of this post are the least important thing I'm giving you.
What does the work is that spec-writer writes for a machine instead of for me, and enforces it with a lint that runs before anything gets created. What does the work is memory holding the things no amount of codebase reading would surface. A dark-mode selector that matches more than it looks like it matches. A notification path that parses JSON with a regex and breaks silently if you change the shape.
The one-paragraph prompt was enough because everything else had already been written down.
Start there. Write one thing down.
The files
Both skills, as they actually run:
compose-workis the orchestrator. It takes a description or an issue URL, runs the spec phase, then creates the issues and queues the agent tasks without stopping to ask anything else.spec-writeris the quality bar. Interview, decomposition, the ambiguity lint above, and the issue templates. It's the longer of the two and the one worth reading closely.
Three clauses are redacted, marked inline where they occur. They described specifics of my own host that don't belong in a blog post. Nothing else is changed.
Take them, break them, tell me what I got wrong.
–Jeremy