We run a fixed set of prompts against every model in the catalog and keep the outputs — so you can compare what they actually produce, not what the docs claim.
Outputs generated 2026-08-26. Costs are measured from the provider's own billing where reported. Nothing here is simulated.