The cost of being polite to your model
A tiny, concrete experiment from this week, and a larger point about prompt hygiene.
I ran a 500‑query benchmark with two versions of the same prompt. One was polite, please, thank you, niceties. The other was terse. Same model, same task. The terse prompt was 7% more accurate and 18% cheaper on tokens. Not a universal rule. A useful data point.
Polite prompts add tokens without adding signal. Tokens are a distraction budget. Every token you spend on tone is a token you did not spend on the task. This is not a case against kindness; it is a case for knowing which parts of your prompt are doing work.
Most production prompts I audit are 3× too long. The first thing I cut is the roleplay intro. The second is the instructions for cases that never occur. The third is the apology clauses. What's left is usually the actual task.
Panio's health assistant got a 40% latency improvement this week from exactly this exercise. Users noticed before the team announced it, which is the gold standard for a performance win.
If you have a prompt you are proud of, reply with it. I am collecting a small gallery for a future Brief.
One email a week. No noise, easy unsubscribe.