Meta / Contents
The Finish Line Keeps Moving
Before TutorPro's soft launch, I had agents rebuild its help center from the source code. They found a tutor who couldn't accept a parent's request while 13,758 tests passed, and every pass since has found more. Quality is still the hard part of agentic development, and I want your ideas for closing the gap.
Here's what I typed into Claude Code on Monday morning:
ALL of that work has been completed by composer. We probably need to /tag-release and get everything in production. We also had some follow up work to do, I think now that everything is done. Can you help get this over the finish line?
By Tuesday morning the finish line had moved four more times.
The work in question is TutorPro, the tutoring platform I built for my wife's business. We're getting it ready for a soft launch. Our own tutors and families move onto it first, before we invite anyone else. It hasn't launched to anyone yet, and after this week I'm glad.
In July I wrote about the factory that builds my software, and near the end I admitted to the gap I hadn't closed: "Autonomous test automation is still the hole in the factory floor." My agents write tests, and CI gates every merge, but nothing in the pipeline drives the app the way a person would.
This post is what falling through that hole looks like. I'm writing it down because I don't think the hole is mine alone, and because I'd like your help closing it.
One rule for the writers
It started last Wednesday with a smaller job.
TutorPro's help center was out of date, screenshots included. I asked Claude to rebuild all of it from the current code, with visitors, parents, and tutors each seeing their own articles. I asked for "a VERY polished help system." I also wrote this, to my credit:
This is a large effort and probably needs a proper plan, before you use a swarm of subagents to do all the work.
So the plan came first. Three agents read the source code behind each audience's screens and wrote an inventory before anyone drafted a word of help.
Then came the rule that turned a documentation project into something else. Every writer got it in their instructions:
If a flow is broken or unreachable, do not write around it. Leave the article out and report the defect with the file and line.
The writing ran as a swarm of 32 agents, with Claude Fable 5.1 coordinating. Sixteen writers each took a slice of the product, and an independent fact-checker followed each writer, checking every claim against the code and the running app.
The screenshots got rebuilt too. The old ones were cropped with hand-typed pixel coordinates, which is why they'd gone stale. The new ones crop to an element on the page, so one command regenerates all of them.
By that evening I had 87 articles, 123 screenshots in light and dark mode, and a pile of reports I hadn't asked for.
Every test green, the journey broken
The inventory had found the worst of it before the writers even started.
On TutorPro, a parent finds a tutor and sends a connection request. The tutor gets a notification and is supposed to accept or decline.
The tutor couldn't. The Accept and Decline buttons only rendered on a tutor dashboard we'd retired, and the notification linked to a page that didn't exist. Tap it, get a 404. The first handshake between a family and a tutor, the one the whole product depends on, went nowhere.
That same evening the test suites ran on the help-center branch. Every one of them passed.
| Suite | Tests | Result |
|---|---|---|
| Unit | 6,332 | Passed |
| Storybook (component renders) | 3,197 | Passed |
| Cloud Functions | 3,869 | Passed |
| Security rules | 360 | Passed |
That's 13,758 green checks, and not one tutor who could say yes.
It wasn't the only one. Most of the notification preferences saved fine, and nothing ever read them. Co-parent management was built and then never mounted anywhere a user could reach it. A "Switch Role" button showed a confirmation toast, reloaded the page, and left you in the same role. An in-app tip told parents they could reschedule a session, and there was no reschedule button.
Even the screenshots caught something. Every dialog in the app had a solid black backdrop instead of a dimmed one, because the move to Tailwind CSS 4 dropped the utility class that did the dimming. Eleven components had it. Nobody noticed until a screenshot showed a thick black band around every modal.
Here's the part I keep turning over. None of those tests were wrong. Each one checked the piece it was written for, and each piece worked. The button rendered. The preference saved. The notification sent. A real person's path runs between the pieces, and nothing in my pipeline walks it.
The fact-checkers made it concrete. Twelve articles ended up held back because the feature each one would describe didn't work. The checkers also logged 217 findings with file-and-line evidence: copy that contradicted the app, links that went nowhere, flows that worked for one role and broke for another. Agents re-verified each one against the code and grouped them into 34 issues.
I asked for help docs. I got an audit.
Why the finish line keeps moving
An audit produces a list, and working through lists is what Composer, my agent factory, is for. On Thursday morning I had Claude queue every issue, ordered so that fixes touching the same files would run one after another.
That went sideways first. Every task Claude created with a dependency came back with a server error, but the task still showed up in the queue, so Claude read each error as a glitch and kept going. It did that about a dozen times. Each error was the Composer server crashing and taking the running agents down with it. Fifteen tasks failed before Claude stopped and told me, "I'm sorry about the mess." The bug was mine, in Composer's task API, and it's fixed now. Still, the verifier had misread its own evidence, and that turned out to be the theme of the week.
By Monday morning the queue was empty and every issue was closed. That's when I typed the finish-line message. Claude's first report back already had a catch in it. Composer's agents had edited 72 help articles along the way, each one updating whatever its own fix touched. Even so, 86 of the 87 articles now cited source files that had changed since a fact-checker last read them.
The line moved four times from there.
- A fix broke a screen. Seeding test data for the held articles turned up a regression. A security fix had done its job, and as a side effect every connection request now showed the tutor "Unknown User" where the parent's name belonged. I held the release.
- The re-check found more. A second swarm re-verified all 87 articles against the new code, and 66 needed corrections. It also found five more data-exposure bugs. I won't describe them until the fixes ship, but every one of them had passed every test. My standing rule now is that any exposure gets filed and queued the moment it's found. Twenty more tasks went into the queue.
- The factory outran the docs. The second pass lived in one big pull request: the held articles plus corrections to dozens of others. While it waited for me to merge it (I forgot), Composer kept merging fixes, and each fix edited the articles it touched, just as the repo's rules tell it to. Every merge put the PR back into conflict. Claude rebased it and corrected 27 more articles, and then another fix landed. "I can't - more conflicts," I typed, and decided to let the whole queue drain before releasing.
- The drain found more. On Tuesday morning, with the queue clear at last, one more pass corrected 26 articles and turned up five more defects. I queued those too.
Ninety-nine pull requests merged into TutorPro between October 1 and Tuesday morning, 52 of them on a single day. The rule that keeps each help article current with its fix is the same rule that kept knocking the big docs PR into conflict.
The passes keep finding fewer problems: 217 findings, then 102, then five, then five more. That's progress. It also means I don't know where zero is.
As I write this, the help-center PR is waiting on a handful of last fixes, and TutorPro's most recent release is from September 27.
It's not just my factory
I'd love to tell you this is a quirk of my setup: one developer, a home-built factory, a lot of agents. I don't think it is, for four reasons.
It's structural. When agents write the code, the tests, and the reviews, the checks inherit the code's assumptions. Composer verifies every task against its own issue, and every fix this week passed its own tests and its own review. Nobody's job was the tutor's whole journey. I drew this shape once before, in Trust, but Verify: a migration bug where every step was locally reasonable and the failure only existed end to end.
Speed outruns proof. There's a name for this now. Werner Vogels called it verification debt in his closing keynote at AWS re:Invent last December, and a post on the CACM blog defines it as "uncertainty about runtime behavior under realistic conditions, not code cleanliness, not maintainability, and not simply missing unit tests." I had 13,758 answers about the pieces and no evidence about the whole.
The numbers agree. Sonar's 2026 State of Code survey found that 96% of developers don't fully trust that AI-generated code is functionally correct, yet only 48% always check it before committing. In Stack Overflow's 2025 developer survey, the top frustration with AI tools, at 66%, was "AI solutions that are almost right, but not quite." Google's 2025 DORA report linked AI adoption to higher delivery throughput and lower delivery stability. Its summary calls AI "the great amplifier."
The people I talk to see it too. Engineers I know who build with agents describe the same pattern. The code arrives faster than anyone can show it works.
Almost right is the expensive kind of wrong, because it passes.
What I'm doing before launch
Some of what happened this week worked, and I'm keeping it.
The docs audit is the big one. Writing for a real reader made every agent walk a flow one step at a time, the way a person would, and the rule against writing around broken flows turned each walk into a test. It's the closest thing to journey testing I've had.
The independent fact-checkers earned their keep too. A writer that believes its own draft won't catch its own mistakes. A second agent, told to check every claim against the code, changed 60 of the first 87 articles.
The rule that every agent changing a screen updates its article stays, even though it caused this week's conflicts. Once the big docs PR lands, it's what keeps the help center from going stale again.
What I'm adding is people.
Back in February I wrote that I was putting TutorPro through real-world testing with my wife, "who happens to be the toughest product tester I know." My wife is first in line again. I'm also walking every role myself, as a visitor, a parent, and a tutor, start to finish, before any family signs in.
The soft launch is the test after that. Our own tutors and families go first, in a small group, because some problems only show up when real people try to get real work done. I'd rather someone we know finds the next "Unknown User."
Now the optimistic part. The factory found all of this before any family did. Agents found, filed, and fixed nearly all of it in six days, and the same speed that creates the problem is what cleared the backlog. I just want to find the next one on purpose instead of by accident.
Help me close the gap
This is where I'd like your help. I don't have good answers to these yet, and I suspect some of you do.
- Journey testing. Who or what walks your app end to end, as each kind of user? Agents driving a browser, scripted scenarios, or people with a checklist? Which one catches the most?
- Builder versus verifier. When agents write the code and the tests, how do you keep verification independent? A different model? Tests written from the spec before the code exists?
- Docs and tests that stay true. When code changes faster than anyone can read it, what keeps your documentation and your test suite describing the app you actually have?
- Release criteria. How do you decide you're ready when every pass finds more? Is it a number, a gate, or a gut call?
Tell me in the comments on LinkedIn or through my contact page. I'll write up what I learn.
The finish line will stop moving eventually. I want it to stop because the checks got better, not because I stopped looking.
–Jeremy