PVL.AI
Open menu

Whitepaper · LLM Cost

The cheap model
costs the most.

List price barely predicts what you actually pay, and says nothing about total cost of ownership. Same prompt, same cap: a 53× spread in actual output tokens.

  • PDF · 8 pages
  • English
  • v1.0 · 2026-09

What this whitepaper covers

01

List price lies

Same prompt, same cap — actual output ranged from 14 to 744 tokens. What to benchmark instead.

02

Match tier to failure cost

Not "how hard is the task" but "how long until a mistake surfaces". Includes our production routing table.

03

max_tokens is a safety valve

Setting it low saves nothing — it turns input you already paid for into scrap. We hit this three times.

04

Caching fails silently

Too short a prefix and the cache is never created — with no error. Parallel requests collide.

05

Two batch misconceptions

The 50% is real, but it does not stack with caching, and it backfires on verify-and-retry work.

06

Observability and two incidents

Provider-reported cost cannot be trusted. Plus 96 calls that all failed for a non-model reason.

Measured in the paper

53×

spread in actual tokens on one prompt

3

times we hit the same max_tokens trap

42pp

terminology gain from a full cached prefix

Who it is for

  • Engineering leads deciding which model handles which task
  • Teams whose bill is climbing without a clear reason
  • Anyone evaluating how much a cheaper model would actually save

Every figure comes from measured production runs, not estimates. Disclosure policy: we name the models we chose, not the ones we rejected — rejection results hold only for our tasks.