A one-person studio with an AI team
I make apps alone: a songbook, an ear trainer, a Czech grammar trainer, a printing utility, a website that finds guitar chord fingerings, and a statics trainer for engineering students. Alone, but not single-handed. Most of the code is now written by AI coding agents, and a second model reviews it. This is how that works in practice, what it changed, and what it did not.
Who does what
- I decide what to build, write the request, read the result and say yes or no. Taste, priorities and anything that touches people (their data, their money, their time) stay with me.
- A coding agent (Claude Code) reads the repository, writes the change, runs the tests, starts the app, takes screenshots and reports what it did and what it did not check.
- A second model as a reviewer (Codex, sometimes Gemini) reads the plan or the diff with fresh eyes. It has no stake in the change being right, and it finds different things.
The agents are fast and tireless, and they are confidently wrong often enough that the whole process is built around catching that.
A rules file that grows from mistakes
Each repository has a file of instructions for agents. It is not a style guide. Almost every line in it records something that once went wrong:
- "Pages and API answers are never cached for long, only
no-cachewith an ETag." Once a one-day cache on the API left a user with a black screen after a deploy: the browser kept an answer in the old format, and the new page crashed on it. - "Never load real ads automatically: not in tests, not in screenshots." Automated views count as invalid traffic, and that can get an advertising account banned.
- "Scripts never kill other processes by name; start a process and stop it by its PID." A pattern-based kill once took down unrelated work.
- "Feedback on the 3D hand is a bug report: reproduce it on the same frets and fingers, and add the shape to the test corpus."
The file is the team's memory. A new session of the agent knows nothing of yesterday; it does know the rules.
Checks that cannot be argued with
An agent can explain why its change is fine. A test cannot be talked into passing. So the rules say what must be true, and scripts check it:
- Type checks and unit tests before anything else.
- Parity with the original. The chord engine exists in Kotlin (in the Android songbook) and in TypeScript (on the website). A test compares the two on data exported from the Kotlin code, down to the order of the shapes.
- A smoke check in a real browser after every deploy: hundreds of checks on the live site, in light and dark themes, on a phone-sized screen, with the API down, with the API answering in an old format, with storage blocked.
- A quality floor. The 3D hand is checked on a corpus of about 400 chord shapes. A change may not add a new kind of problem, make any problem more frequent, or move the fingers of good shapes. The floor may only go up.
- Bit-for-bit refactoring. When the agent speeds up the pose solver, the answers for the whole corpus must match the old ones byte for byte. Faster and different is a new solver, and a new solver needs a new review.
A second model reviews the first
Before a large change I ask a second model to review the plan; after it, the diff. It is the cheapest check there is, and it finds real things. On the day I write this:
- Reviewing new pages for this site, it pointed out that the legal footer should not claim "not a VAT payer": that is not required, and it can go stale.
- Reviewing a privacy policy against the app's code, it found that the songbook sent song titles to analytics and showed ads in Europe without asking for consent first. Both were fixed in the app, not just in the policy.
- Reviewing that fix, it found that a screen rotation could show the consent form from an activity that no longer existed.
The reviewer is wrong sometimes too. The rule is not "do what it says", it is "every point gets an answer": fixed, or explained.
Where a person still has to look
Tests check what we thought of. Eyes catch the rest:
- Screenshots are looked at, not just taken. Tests do not notice notes overlapping on a staff, a label running off a phone screen, or a page that is technically present but black on a dark background. Every visual change ends with screenshots sent to my phone.
- Sound is listened to. Nothing headless can hear a wrong chord.
- Claims are checked. An agent once reported a check as done that it could not actually have done (a consent dialog "outside Europe" on a machine that is in Europe). It noticed, said so, and corrected itself, which is exactly the behaviour I want. But it is why "verified" in a report is a claim to check, not a fact.
What changed
- Scope. One person can keep several apps alive and improving: interfaces in up to seven languages, accessibility, privacy reviews, store listings, tests for all of it.
- Speed of the boring parts. A migration, a release with screenshots, a policy brought in line with the code: hours, not weekends.
- Quality became explicit. When work is delegated, "good" has to be written down. Writing it down made it better than when it lived in my head.
What did not change
- Deciding what is worth building.
- Responsibility. The agents write the code; I ship it, and it is my name on the store page and in the privacy policy.
- The need to understand the code. I review what goes out. An agent is a strong colleague, not a reason to stop knowing how things work.
If you want to see the result: the apps and the chord site are all built this way.