Shipping the same day a runner asks: how I build with Claude Code

My loop with Claude Code on two AI products: sessions that message each other, research that argues against itself, and testing done before I look.

My Claude Code sessions message each other. They test my changes before I look, and they re-film my demo videos when the app changes. My median message to them is under 20 words.

I spent the summer as an AI-native builder on two products: Fluentide, an AI-powered Chinese listening app with 2,000+ learners, and TAOSport, an AI running coach with 250+ active runners. Twenty words are enough because Claude runs a loop that starts and ends with the people using the app. Here's that loop.

A loop that starts and ends with users: Hear, Decide, Build, Check, and Back to the user, around Fluentide (2,000+ learners) and TAOSport (250+ active runners).

Hear

Requests from people reach me in two ways. After a team meeting I paste the raw transcript: four agents read the code in parallel, and Claude comes back with a plan for every request in it. Bugs come from runners, as screenshots in our group chat, and a screenshot plus a few words is enough for Claude to reproduce the problem on its own.

On Fluentide, though, more of the work started from our data than from anyone's message. That loop starts in a different place, so it gets its own post, the next in this series. This one follows the requests people bring.

Decide

A plan is only as good as the research behind it, and Claude's research comes back confident and well cited. So I make it argue against itself: it studies how others solve the problem, tries to prove each finding wrong with our own code, history and usage, and only then advises.

That step often changes the plan. When Claude studied Telegram's code and 20 other chat apps for our coach's chat, it named our biggest problem, citing about ten past bug fixes. Made to prove itself wrong against our own code, it recounted: one. Six of its recommendations, including the one it ranked first, were overturned or scaled back before the plan was written.

Most decisions are smaller than that, and they shouldn't wait for me. Agents like to end a task with a list of open questions; mine settle the small calls with evidence and write the call and its reason into the commit. Only four kinds come back to me: logins, money, taste calls I keep for myself, and anything public or irreversible.

Build

Claude Code sessions can message each other. It's called cross-session messaging, and many people don't know it exists. So once a plan is settled, I give each step its own session and let them coordinate. On one plan they agreed who owned which files, confirmed "226 tests pass on the merged file", and passed findings across. I didn't relay a single message. The most useful bug report came from the session next door: it caught our AI grader failing correct translations, and told the session testing the grader.

Three real messages between Claude Code sessions on one Fluentide plan: one confirms 226 tests pass on a merged file, one proposes which files each session owns, and two minutes later the other confirms the split and reports that the AI grader's verdicts argue themselves out of the flag.

When an app changes, its demo video usually goes stale. Mine don't, because Claude is also my video editor: it drives the real app on camera and times every cut from the measured narration, never by hand. One command re-films the demo and re-cuts the edit in about two minutes. One demo survived four rewrites of the app it was filming without me touching a single timing; its captions still light up word by word, and a click still lands exactly on the spoken word "back".

Check

A change isn't done when the code is written, but for most changes I don't click through anything myself. Claude runs the change end to end in Ego Lite, a browser built for agents, logged in like a real user, and fixes what it finds before reporting. Then it hands me a storyboard: one captioned screenshot per step, loading, empty and error states included. Scrolling through it is my review.

Six frames of a storyboard from TAOSport's AI chat, translated from Chinese: an unanswered question with Retry, a failed retry, the answer, thumbs-down reasons, a table streaming in, and the switch-chat dialog.

Changes to our AI coach need one more check, since a screen can look right while the coach's answer is wrong. Claude replays the exact words the runner complained about in the real app, saves the answer with screenshots and a full trace, then reads it and signs a verdict. The fix counts only when that runner's own question gets a good answer.

Back to the user

That runner is also where the loop ends. It closes in the chat where it started: every change gets posted back to the group, and many requests shipped the same day a runner asked.

That is why my messages can stay short. Claude does most of the building and checking. My time goes to the people in that chat, and to the calls only I can make.