List price lies
Same prompt, same cap — actual output ranged from 14 to 744 tokens. What to benchmark instead.
Whitepaper · LLM Cost
List price barely predicts what you actually pay, and says nothing about total cost of ownership. Same prompt, same cap: a 53× spread in actual output tokens.
Same prompt, same cap — actual output ranged from 14 to 744 tokens. What to benchmark instead.
Not "how hard is the task" but "how long until a mistake surfaces". Includes our production routing table.
Setting it low saves nothing — it turns input you already paid for into scrap. We hit this three times.
Too short a prefix and the cache is never created — with no error. Parallel requests collide.
The 50% is real, but it does not stack with caching, and it backfires on verify-and-retry work.
Provider-reported cost cannot be trusted. Plus 96 calls that all failed for a non-model reason.
53×
spread in actual tokens on one prompt
3
times we hit the same max_tokens trap
42pp
terminology gain from a full cached prefix
Every figure comes from measured production runs, not estimates. Disclosure policy: we name the models we chose, not the ones we rejected — rejection results hold only for our tasks.