“Return bounding boxes for each shape as JSON: [{"label":"...","box":[x_min,y_min,x_max,y_max]}] in pixel coordinates. Image is 400x220. JSON only.”
“coords-within-tolerance”
imagetext → text 200K context
Each capability below is a fixed experiment: the same prompt and the same input images for every model we test. That's what makes these outputs comparable — and it's why the same bicycle shows up on every image model's page.
Each of these is a fixed experiment — the same prompt and the same input images for every model we test, so the outputs are directly comparable.
“Return bounding boxes for each shape as JSON: [{"label":"...","box":[x_min,y_min,x_max,y_max]}] in pixel coordinates. Image is 400x220. JSON only.”
“coords-within-tolerance”
“In this bar chart, which bar is tallest: left, middle, or right? One word.”
“Right.”
“How many shapes are in this image in total? Reply with just the number.”
“got=5 want=5”
“Describe exactly what you see in this image: the shapes and their colors.”
“The image contains two geometric shapes: 1. A red square. 2. A blue circle. The red square fills the left side of the ”
“What text appears in this image? Reply with the text only.”
“HIFIBOTS”
“Is the square to the left or to the right of the circle? One word.”
“Left”
“What color is the square in this image? One word.”
“Red.”
The same model, priced by each provider that serves it. Cheapest first — the card's headline price is the best available anywhere, which is the wrong number for choosing where to actually send your traffic.
| Provider | Price | Context | |
|---|---|---|---|
| openrouter | $0.25 / $1.25 per 1M | 200K | Get a key ↗ |
| anthropic | Not published | 200K | Get a key ↗ |
Every claim is tagged by how we know it. Tested means we ran the model against a fixed prompt and checked the output — the images above are those runs. Researched means it's documented by the provider and cited. Nothing here is inferred.