Batch workload benchmark for 300,000 records with 150 input tokens and 150 output tokens per record, balancing blended spend against latency, tail risk, and output-heavy pricing.
Sweet spot:
Pricing structure:
Long-tail spread:
| Rank | Model | Provider | Blended Cost | Median Latency | Median Batch Time | Flag |
|---|
The highlighted rows mark the current Pareto frontier, where no other benchmarked model is both cheaper and faster at the same time.
The stacked cost chart shows that output pricing quickly becomes the larger share of spend even when input rates look favorable, which is why the misleading flag matters for this equal-token workload.
The batch window chart translates per-request latency into a more operational question: how long the full 300,000-record job could take under median and p99 conditions.