Sixty-One Percent More Thought, Five Percent More Prose

Opus 5 and Opus 4.8, Measured on 1,097 German Sentences

Anthropic released Opus 5 on July 24, 2026. I had, at that moment, a job queued that wanted exactly this sort of model. I have been expanding the verb corpus of Konjugieren, my German-conjugation app, from the 990 verbs it shipped with to 3,572, and example sentences for most of the new arrivals came out of a corpus of real German text. For 1,097 of them the corpus had nothing, because they are too rare to appear in one that would fit comfortably on my SSD, so somebody was going to have to write 1,097 example sentences from scratch. The timing offered something better than a benchmark, because a benchmark measures a model on a task chosen for being measurable, and I had a task I actually needed done. So I split the work down the middle, gave half to the outgoing Opus 4.8 and half to the incoming Opus 5, held every other variable I could think of still, and instrumented all of it. I was curious about four unglamorous things: time, token use, cost, and verbosity.

Continue reading

The Aside That Built a Test Suite

Two Asides, Nine Stale Claims, and a Knight Named for a Bear

An emoticon-embellished joke that I made to Claude Code led me down a deep rabbit hole. This joke, seven words long, was not a request for the coding agent to perform work. But by the end of the day, the joke had resulted in a new tool in my repository, fixes for nine documentation defects, the discovery that one of my app’s features was 72 percent unimplemented, and a conversation about the murder of Thomas Becket that taught me something about Germanic bear taboos.

Continue reading

The Fable and the Opus

Anthropic’s New Model, Measured as a Code Reviewer

A fable is a story that speaks: Latin fabula, from fari, “to speak”, and in the genre as Aesop and La Fontaine practiced it, the speaking is done by animals. An opus is a work: the crafted thing itself, named for its workmanship. Today, June 9, 2026, Anthropic released Fable 5, a new model positioned above Opus 4.8 and priced at twice Opus’s rate per token. I wanted to know, before release day ended, what that premium buys a working iOS developer. So I staged a contest of genres: four Claude Code sessions, two models at two effort levels apiece, each session given an identical request to review the codebase of Conjuguer, my French-conjugation app, and to rank what it found by impact. This post is the comparison, with tables. Being about a fable, it ends with a moral.

Continue reading