Multimodal stopped being a feature
Every release this month treats image and video input as table stakes. That quietly resets what a baseline product looks like.
I went back through the releases of the last six weeks and could not find a significant one that shipped text-only. Image understanding is now assumed, video is common, and audio is close behind. The product consequence is bigger than the capability consequence. Users who have used one multimodal assistant now expect every assistant to accept a photograph, and a text box on its own reads as dated in a way it did not six months ago. At Panio this is the whole product: someone photographs a vet report and the app turns it into something trackable. Two years ago that was the hard part. It is now the easy part.
ByteDance released Seed 2.1 Turbo this week and DeepSeek put out a Flash variant at the end of July. Both are aimed at the same target: high throughput at low cost for workloads that do not need deep reasoning. That is most workloads. If you are routing everything to a frontier model out of caution, you are paying a premium for latency you could be spending on evaluation.
Enforcement is a week old. Nothing public yet, which is what I expected. Regulators move slowly at the start and then in clusters. The useful thing to do with this window is documentation, not lobbying.
Pericls now runs an expansion simulation: pick a product and a candidate market and it returns the obligation delta before you commit engineering time. It is the feature I most wanted to exist when I was doing payments privacy work at Google and had to answer the same question by hand, over weeks.
One email a week. No noise, easy unsubscribe.