The week efficiency out-argued capability
Grok 4.5 competes on tokens, not IQ. Full-duplex voice ships. And Cowork's usage data shows where agents actually earn their keep.
xAI shipped Grok 4.5 and did not lead with a benchmark. Musk called it 'Opus-class but much faster' and priced it at $2 / $6 per million tokens, under both Opus 4.8 and GPT-5.6 Sol, claiming roughly 2x token efficiency: about 15,954 output tokens to finish a SWE-Bench Pro task where Opus 4.8 spends around 67,020. Treat the numbers as xAI-reported until someone independent checks them. The framing is the real story. The competitive axis just moved from 'who is smartest' to 'who finishes the task for the fewest tokens,' and cost-per-completed-task is the one model metric that shows up on a P&L. Mistral made the same point from the other end, shipping Leanstral 1.5 as open weights: a 119B / 6B-active model, Apache 2.0, that reports 100% on the miniF2F formal-math set. Not state of the art, saturated. When a narrow open-weight model runs the table on a structured domain and you can self-host it, the build-versus-buy question for that vertical is closed.
OpenAI released GPT-Live-1, a full-duplex voice model: you can interrupt it, it can listen and speak at the same time, and it hands the hard reasoning back to GPT-5.5. Full-duplex is the line between a voice agent that feels real-time and one that feels like a walkie-talkie. There is no API and no pricing yet, so this is a capability note, not a roadmap input. But if you are building anything conversational, this is the bar your users will start measuring you against.
Anthropic pushed Claude Cowork to web and mobile and, more usefully, published what people do with it. About a third of sessions are 'pull scattered updates into one report' and 'reconcile these spreadsheets.' Sixteen percent is content. Nine percent is software development. Knowledge work, not coding, is now roughly half of agent sessions. The beachhead for agents was never the IDE. It is the boring reconciliation work nobody wants to do and everybody has to.
The EU AI Act's enforcement powers over general-purpose model providers go live August 2, with fines up to 3% of global turnover or 15 million euros, and the Article 50 transparency duties, chatbot disclosure and AI-content labeling, arrive with them. Signing the GPAI Code of Practice buys a presumption of conformity. If you serve EU users, the posture you want on August 2 is one you decide this month, not that week.
Spent the week rebuilding a client's model-routing table after the new pricing landed, and the frontier tier shrank. Half the calls I had parked on a flagship now go to a workhorse model at a tenth of the cost with no measurable quality loss on their own eval set. That is the whole exercise: not 'which model is best,' but 'which is the cheapest model that still passes the test.' If you have a routing table you have not re-run since spring, reply and I will send the template I use.
One email a week. No noise, easy unsubscribe.