LLM Cost Benchmark Dashboard

Batch workload benchmark for 300,000 records with 150 input tokens and 150 output tokens per record, balancing blended spend against latency, tail risk, and output-heavy pricing.

Models Benchmarked
Lowest Blended Cost
Fastest Median Latency
Flagged as Misleading
Decision framing
Bubble size = context window Diamond markers = misleading unit-price signal Log-scaled cost axes preserve low-end detail Median and p99 included for batch SLA context

Sweet spot:

Pricing structure:

Long-tail spread:

Frontier models
Models that are not beaten on both cost and median latency at once.
Median benchmark point
Useful reference line for separating the lower-left cost/latency quadrant.
Highest blended run
Upper bound in the current benchmark set for this fixed batch workload.
Blended cost vs. median latency sweet spot
Total blended cost by model
Provider median benchmark profile
Input vs. output cost mix for the lowest-cost contenders
Median vs. p99 batch completion time for the cheapest shortlist
Top 10 lowest blended-cost models
Rank Model Provider Blended Cost Median Latency Median Batch Time Flag
Shortlist reading notes

The highlighted rows mark the current Pareto frontier, where no other benchmarked model is both cheaper and faster at the same time.

The stacked cost chart shows that output pricing quickly becomes the larger share of spend even when input rates look favorable, which is why the misleading flag matters for this equal-token workload.

The batch window chart translates per-request latency into a more operational question: how long the full 300,000-record job could take under median and p99 conditions.