← The Brief
46 ·

On evals nobody writes

Three evals I wish every AI team was running, and none of them are on a leaderboard.

The eval nobody writes

The 'boring cases' eval. Not the adversarial, jailbreak, red‑team set. The 200 normal questions a normal user asks on a normal Tuesday. Most teams can tell me their MMLU score. Almost none can tell me the accuracy on their own top‑200 queries. Fix that first.

A pattern I keep seeing

Teams confuse 'model is better' with 'product is better'. Put a human in the loop, measure task completion, and you will find that a worse model with better tooling beats a better model with worse tooling roughly eight times out of ten.

One number

The quiet metric I track on every agent project: percentage of user sessions that end without the user undoing something. It is a crueller metric than satisfaction, faster than retention, and impossible to fake.

One thing I'm building

Pericls shipped a new horizon‑scanning view this week that surfaces regulation changes by business impact, not by jurisdiction. It is the first view that made me want to open the product at 7am, which is how I know it is working.

A quote I came back to

From an old Andy Grove memo: 'When everybody knows that something is so, it means that nobody knows nothing.' The AI industry has a lot of everybody‑knows right now. Keep a list. Revisit it in a year.

Get the next one

One email a week. No noise, easy unsubscribe.

← Nº 45
What I'm telling portfolio CTOs this month
47
The week agents started paying for things