On evals nobody writes
Three evals I wish every AI team was running, and none of them are on a leaderboard.
The 'boring cases' eval. Not the adversarial, jailbreak, red‑team set. The 200 normal questions a normal user asks on a normal Tuesday. Most teams can tell me their MMLU score. Almost none can tell me the accuracy on their own top‑200 queries. Fix that first.
Teams confuse 'model is better' with 'product is better'. Put a human in the loop, measure task completion, and you will find that a worse model with better tooling beats a better model with worse tooling roughly eight times out of ten.
The quiet metric I track on every agent project: percentage of user sessions that end without the user undoing something. It is a crueller metric than satisfaction, faster than retention, and impossible to fake.
Pericls shipped a new horizon‑scanning view this week that surfaces regulation changes by business impact, not by jurisdiction. It is the first view that made me want to open the product at 7am, which is how I know it is working.
From an old Andy Grove memo: 'When everybody knows that something is so, it means that nobody knows nothing.' The AI industry has a lot of everybody‑knows right now. Keep a list. Revisit it in a year.
One email a week. No noise, easy unsubscribe.