2 min lesson
Artifacts: the feature that made the shift believable
Explain your answer to "Why did artifacts (proof-of-work recordings), rather than a smarter model, drive adoption of async cloud agents?" Add one concrete detail from the lesson.
Step 1 of 2
Artifacts: the feature that made the shift believableproof of work, not lines of code
The decision that unlocked async adoption wasn't a smarter model - it was proof of work. Artifacts gave a cloud agent the ability to navigate its own machine and come back with evidence: for a front-end change it spins up a browser, tests the change, records a video and links the recording to its message. The human reviews an AI's work the way they'd review a junior engineer's - by the result, not by reading every diff.
Releasing artifacts (around early January) caused a measurable uptick in cloud-agent adoption and, internally, pushed cloud agents to roughly 30-40% of PRs created or merged. That's the move a PM should be able to narrate: the constraint was reviewer trust in unwatched work, the lever was proof-of-work artifacts, and the metric that confirmed it was share-of-merges, not a vanity count. Output became outcome, and the adoption number followed.
When the room asks for a product opinion, resist the matrix and tell this as a decision: “The third era moves attention from output to outcome, which makes review - not generation - the bottleneck. So the highest-value bets build reviewer trust. Artifacts did exactly that: proof-of-work recordings let humans review the result, and adoption jumped to ~30-40% of cloud-agent PRs.” Then name your own next bet against the same bottleneck and pre-commit its kill criterion. A narrated decision with a number reads as judgment; a framework reads as prep.