🔬 Tested by TokenRouter

Same prompt. Every model. Judge for yourself.

We run a fixed set of prompts against every model in the catalog and keep the outputs — so you can compare what they actually produce, not what the docs claim.

The prompt Edit this image: change the red square to green. Keep everything else identical.

Outputs generated 2026-07-17. Costs are measured from the provider's own billing where reported. Nothing here is simulated.