Picking an image model family without trusting a leaderboard
Updated 2026-10-02
Image model rankings go stale in weeks, and a leaderboard average says little about your use case: product shots, poster text, character consistency and photo edits stress different things. This page gives a decision framework based on what you can verify, then a method for settling the question with your own prompts through one endpoint.
The families in the catalogue
The catalogue spans several model creators. The model families most people shortlist are:
- GPT Image (OpenAI):
gpt-image-1,gpt-image-1.5,gpt-image-2and related variants. - FLUX (Black Forest Labs): FLUX.2 variants such as dev, pro, max and flex, Kontext editing models, and
flux-3-image. - Seedream (ByteDance): Seedream 4.0, 4.5 and 5.0 variants, including Flash, Lite and Pro.
- Nano Banana (Google): Nano Banana 2 and Pro, along with Gemini image models that work through the same endpoint.
- Qwen-Image (Alibaba), Ideogram, Grok Imagine, Recraft and others.
Browse the models index to see exactly what is live; ids change as models are added.
Decide by constraint first, quality second
Most selection problems are settled by hard constraints before any visual judgement:
| Question | Why it narrows the field |
|---|---|
| Do you need to edit existing images? | Only edit-capable models accept input_references. Drop the rest. |
| Do you need to control output size? | Some models honour resolution, aspect_ratio and size; others generate at a fixed size. |
| Do you need OpenAI-only controls? | quality, background, output_format and input_fidelity apply to OpenAI models only. |
| How many references must you pass? | Reference caps differ; for example the platform documents up to 10 references for flux-3-image. |
| What is your volume and per-image budget? | Per-image price differs a lot between and within families; see the live table. |
| Do you need text in the image? | Rendering of legible text varies by model; test it explicitly rather than assuming. |
After constraints, the remaining differences are largely qualitative: aesthetic defaults, how literally the model follows long prompts, how well it keeps a subject consistent across edits. I will not rank these here, because any ranking not backed by your prompts is a guess.
A tiering pattern that holds up
- Draft tier: a fast, low-cost model for exploration, thumbnails and internal review.
- Final tier: a premium model for assets that ship.
- Edit tier: an edit-capable model for revisions of approved outputs.
The same endpoint serves all three, so the tiers are a config map from role to model id, not three integrations.
| Model | Cheapest host | Priciest host | Cheapest is | Hosts |
|---|---|---|---|---|
| black-forest-labs/flux.2-dev | MachGen $0.0031 / image | Fal $0.012 / image | 74% lower | 2 |
| google/nano-banana-2 | MachGen $0.034 / image | Fal $0.08 / image | 57% lower | 2 |
| google/nano-banana-pro | MachGen $0.0672 / image | Fal $0.15 / image | 55% lower | 2 |
| black-forest-labs/flux.2-pro | DeepInfra $0.015 / image | Black Forest Labs $0.03 / image | 50% lower | 3 |
| black-forest-labs/FLUX.1-dev | DeepInfra $0.009 / image | SiliconFlow $0.014 / image | 36% lower | 2 |
| alibaba/wan-2.6 | Atlas Cloud $0.021 / image | DeepInfra $0.03 / image | 30% lower | 2 |
| black-forest-labs/flux.2-max | Black Forest Labs $0.07 / image | DeepInfra $0.1 / image | 30% lower | 3 |
| seedream-5.0-pro | Atlas Cloud $0.036 / image | WaveSpeedAI $0.045 / image | 20% lower | 3 |
| recraft-4.1 | WaveSpeedAI $0.04 / image | Pika $0.042 / image | 5% lower | 2 |
| bytedance/seedream-4.0 | WaveSpeedAI $0.027 / image | DeepInfra $0.04 / image | 33% lower | 4 |
| bytedance/seedream-5.0-lite | Atlas Cloud $0.0315 / image | WaveSpeedAI $0.035 / image | 10% lower | 4 |
| qwen-image-max | WaveSpeedAI $0.07 / image | DeepInfra $0.075 / image | 7% lower | 2 |
Per image, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.
What to write down about each model
As you test, keep a short record per candidate: the exact id, whether it accepts references and how many, whether it honours size controls, what the cost basis is, and any prompt style it prefers. That table becomes your routing config and saves the next engineer repeating the experiment. Note failures as well as successes: a model that returns a refusal or an error on a class of prompts you need is ruled out as quickly as one with poor output.
How to A/B models on your own prompts
Because every model is called the same way, a comparison harness is a loop over model ids. Pull the exact ids from each model's page.
import requests, base64, pathlib
BASE = "https://videorouter.sh/api/v1"
H = {"Authorization": "Bearer llmr_sk_live_..."}
MODELS = ["gpt-image-2", "flux-3-image", "qwen-image-3.0"] # use ids from the model pages
PROMPTS = [l.strip() for l in open("prompts.txt") if l.strip()]
out = pathlib.Path("ab"); out.mkdir(exist_ok=True)
rows = []
for p_i, prompt in enumerate(PROMPTS):
for model in MODELS:
r = requests.post(BASE + "/images", headers=H, timeout=120,
json={"model": model, "prompt": prompt}).json()
if "error" in r:
rows.append((p_i, model, None, r["error"]["message"])); continue
d = r["data"][0]
img = base64.b64decode(d["b64_json"]) if d.get("b64_json") else requests.get(d["url"]).content
f = out / f"{p_i}_{model.replace('/', '_')}.png"
f.write_bytes(img)
rows.append((p_i, model, r["usage"]["cost"], str(f)))
Then review the outputs blind. Shuffle the file names or put them in a spreadsheet without the model column, score each against a short rubric (prompt adherence, text legibility, subject fidelity, artifacts), and join the scores back to the model and usage.cost. The number to compare is cost per accepted image: price divided by the fraction of outputs you would actually ship. A cheaper model with a lower acceptance rate is often not cheaper.
Practical tips for a fair test
- Use 20 to 50 real prompts from your product, including the hard ones.
- Generate more than one sample per prompt for the models you shortlist; single samples overstate differences.
- Keep size and aspect ratio equal across models where they are honoured, and note models that ignore them.
- Test the editing path separately with your own images; see text-to-image vs image-to-image.
- Re-run the test when a new model appears. It costs one list entry.
Hosts also matter: the same model can be served by more than one host at different prices, so check the price spread before pinning one. Cost levers are covered in image generation cost control. Create a key and run the harness above; the quickstart has the base request.
Frequently asked questions
Which image model is best overall?
There is no single best. Results depend on your prompts, whether you need text rendering or edits, and your budget. Compare candidates on your own prompts through the same endpoint.
How do I compare GPT Image, FLUX, Seedream and Nano Banana fairly?
Run the same prompt set through each model by changing only the model id, review outputs blind against a rubric, and compare cost per accepted image rather than list price.
Can I use one integration for all of them?
Yes. They share POST /v1/images. Differences are in supported parameters, such as OpenAI-only quality controls and which models accept input references.
Does the model I pick change which parameters work?
Yes. quality, background and input_fidelity apply only to OpenAI models, and some models ignore size controls.
Keep reading
- Choosing an Image Generation API — Price, Quality and Control
- Text-to-Image vs Image-to-Image API: Modes, Fields, Pitfalls
- Image Generation API Cost Control: Billing, Drafts, Caching
- GPT Image API Parameters: Sizes, Quality, Edits and Output
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →