This is how the story will feel to read.
When code is cheap, engineer to behavior.
scroll plz ↓
I had never shipped a native app. Until this month.
I've spent almost two decades building beautiful and joyful user interfaces for the web. In an internship before college, building a CRM for an underwater robotics company, I discovered that Perl and JavaScript could make powerful apps that ran anywhere there was a browser. That set the tone of my career: founding front-end engineer at Patch, then founding engineer at Kickstarter, Jukely, Abacus, and Lemontree, doubling down on JavaScript each time.
Meanwhile the phone grew up next to me. As far back as 2009, someone at every Patch investor meeting asked where the iPhone app was. I tried to learn Objective-C, then Swift, but something more urgent always came up in the lane I knew. So I got defensive: the web needs no install, works everywhere, and answers to no corporation's app store.
AI has made creating software more accessible. My 11-year-old is building a web app that tracks how much water she is drinking by growing celebrity hair as she drinks. She can now build apps, and scaling them with wise architecture is what she will learn next. I built all of tale.fyi with AI as the hands and my architecture and taste as the plan. When syntax stops being the bottleneck, judgment is what's left.
An e-reader wants to be native: better scrolling, a spot on the home screen. But I don't know Swift, and I can't afford to pay someone who does. At a previous job we considered React Native, which shares code with the web. For tale, I wanted a truly native app, and code is now cheap, so I decided to share behavior instead.
If I could describe the product as platform-independent pass/fail gates, an agent could build a client in any language by iterating until the gates passed. It turns out that I'd been relying on this idea my whole career: every JavaScript engine is measured against a shared suite called test262, and every browser against web-platform-tests. It's called conformance testing.
So I started tale-spec, which I'm open sourcing today. It states the product's rules in prose (your progress is a high-water mark, not a cursor, so scrolling up never makes you less finished), backs the pure logic with JSON test vectors, and describes reader sessions as scenarios: open a book, read, leave, come back. Anything that makes up the product lives in the spec; anything that should feel like the device belongs to the platform. The gates pass on the web from day one, fail on a native app that doesn't exist yet, and the agents build until they pass.
The first decision came on day two: render a whole novel with SwiftUI's lazy list, or with TextKit 2, Apple's lower-level text engine? I had never used either, so I didn't have an opinion. I had to form one. So the agents built both, measured dropped frames and layout time with the same probe, and I side-loaded both onto my phone and scrolled. The numbers and my own eyes agreed on TextKit 2. I had reached my opinion, and the decision was made.
Within a day, the Swift passed all 220 shared vectors, which felt good, but when I opened the app, something wasn't right. Passing only showed that the Swift math matched the JavaScript math, not that someone who reads, quits the app, and reopens it lands where they left off. It was uncanny valley: the high-water read-to marks didn't match the web, and an audio play button was the wrong shape, a dissonance the tests couldn't see.
So I redefined conformance as black-box testing of the real product: a real browser for the web, and the real app on iOS, force-quit and relaunched to make sure your place carries across sessions. Five days later I noticed another flaw in my strategy: prose scenarios let each platform's test adapter interpret them, and an agent will sometimes interpret generously in its own favor. Now the scenarios are executable plans judged centrally by tale-spec; an adapter only turns "the reader opens the book" into a tap or a click. There are 23 of them, and tale-spec does not need to know, or care, what language the client is written in.
The Mac was mostly product decisions, about pointers and menus, and Apple's Catalyst did the rest. Since the scenarios describe actions rather than gestures, the Mac adapter just slots in. The hard part was the disruption of the AI building on my machine: the tests kept stealing my screen and failing whenever the Mac locked, so we moved them into a fresh virtual machine each run. That had a happy side effect of making each run repeatable and isolated.
I typed none of the code. Claude generated most of it, Codex some, and each reviewed the other. What I did was the structure. The ticket defines scope, the spec defines behavior, the web app is the reference, and when they disagree the agent stops and asks. Nothing is done until a gate fails without it and passes with it. A branch that grows past one and a half times its first-review size gets re-cut, not patched, because of one hard evening I had fighting an agent that kept adding more and more slop. And one model reviews every diff another generated, which on the web side caught a plausible database fix that would have crashed weekly, before it shipped. It's how you'd run a team of fast, confident, occasionally wrong junior engineers.
Five weeks from the first spec commit on August 15 to both App Stores on September 18: about 43,000 lines of Swift I didn't type, and an app that follows the same rules as the website. The $200 in AI subscriptions ($100 each for Claude and Codex) covered all of tale in that time, not just the native app. Now I can open tale from my home screen on the subway, without a connection, and listen to a few paragraphs of The Great Gatsby.
No Swift experience required, but eighteen years of architecting software certainly helped me set the shape that the agents lived in. The phone was the medium I spent nearly two decades telling people I didn't need. When syntax got cheap, the judgment I'd been building on the web could go anywhere.