Keeping per-image spend predictable at volume
Updated 2026-10-02
Image generation looks cheap per call and gets expensive through multiplication: retries, variants, regenerations, and a few high-resolution requests hidden in a stream of cheap ones. Cost control is mostly bookkeeping plus a handful of levers. This page covers how the billing works and how to use each lever.
How billing works
- Billing is per image, and the response carries a
usage.costfield with what the call cost, including the platform fee. Log it on every request; it is your ground truth for cost per image without waiting for an invoice. - The pricing basis differs by model. Some are flat per size and quality combination, some are priced per token (
gpt-image-2, for example), some per output megapixel (the Ideogram and FLUX.2 [dev] rows on Fal), and some per resolution tier. So the sameresolutionrequest can change the price on one model and do nothing on another. The model page and the live table show the basis for each. - A request that fails upstream is not billed.
- Requesting
nimages generates and bills multiple images in one call.
| Model | Cheapest host | Priciest host | Cheapest is | Hosts |
|---|---|---|---|---|
| black-forest-labs/flux.2-dev | MachGen $0.0031 / image | Fal $0.012 / image | 74% lower | 2 |
| google/nano-banana-2 | MachGen $0.034 / image | Fal $0.08 / image | 57% lower | 2 |
| google/nano-banana-pro | MachGen $0.0672 / image | Fal $0.15 / image | 55% lower | 2 |
| black-forest-labs/flux.2-pro | DeepInfra $0.015 / image | Black Forest Labs $0.03 / image | 50% lower | 3 |
| black-forest-labs/FLUX.1-dev | DeepInfra $0.009 / image | SiliconFlow $0.014 / image | 36% lower | 2 |
| alibaba/wan-2.6 | Atlas Cloud $0.021 / image | DeepInfra $0.03 / image | 30% lower | 2 |
| black-forest-labs/flux.2-max | Black Forest Labs $0.07 / image | DeepInfra $0.1 / image | 30% lower | 3 |
| seedream-5.0-pro | Atlas Cloud $0.036 / image | WaveSpeedAI $0.045 / image | 20% lower | 3 |
| recraft-4.1 | WaveSpeedAI $0.04 / image | Pika $0.042 / image | 5% lower | 2 |
| bytedance/seedream-4.0 | WaveSpeedAI $0.027 / image | DeepInfra $0.04 / image | 33% lower | 4 |
| bytedance/seedream-5.0-lite | Atlas Cloud $0.0315 / image | WaveSpeedAI $0.035 / image | 10% lower | 4 |
| qwen-image-max | WaveSpeedAI $0.07 / image | DeepInfra $0.075 / image | 7% lower | 2 |
Per image, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.
Lever 1: resolution and quality
Where a model honours size controls, the price often rises with resolution. Generate at the smallest size that satisfies the display target, and request higher resolution only for final assets. For OpenAI models the quality setting is an additional lever, applied only to those models. Check on the model page whether size actually changes the price, because on models that ignore size controls, asking for less saves nothing.
Lever 2: draft, then final
Run exploration on a cheap, fast model and promote only approved prompts to a premium model. Because the endpoint is shared, promotion changes only the model string. A simple workflow:
- Generate 4 draft variants on the draft model.
- A person or automated check picks one.
- Re-run that prompt (and seed-like inputs) once on the final model at final resolution.
This turns the expensive model's cost into a function of accepted images rather than attempted ones. For edits, use the approved draft as the reference image on an edit-capable model; see image modes.
Lever 3: batching and concurrency
For bulk jobs, queue the work and cap concurrency per model. This does not change the unit price, but it bounds how fast you can spend and keeps you under per-key rate limits (a 429 includes Retry-After). Retry only transient failures (429, 5xx) with backoff, and never retry a 400, which will fail identically every time. If your volume is large and latency-insensitive, check the platform's Batch documentation for whether a discounted route covers your model, since discounts are not available on every model.
Lever 4: cache outputs
The cheapest image is the one you do not regenerate. Cache by a hash of everything that determines the result:
import hashlib, json
def cache_key(model, prompt, size, refs):
blob = json.dumps({"m": model, "p": prompt, "s": size, "r": sorted(refs)}, sort_keys=True)
return hashlib.sha256(blob.encode()).hexdigest()
def get_image(db, store, payload):
k = cache_key(payload["model"], payload["prompt"],
payload.get("size") or (payload.get("aspect_ratio"), payload.get("resolution")),
[r["image_url"]["url"] for r in payload.get("input_references", [])])
if hit := db.get(k):
return store.url(hit)
resp = call_images_api(payload) # your POST /v1/images wrapper
key = store.save(resp) # copy to your own storage
db.set(k, key)
return store.url(key)
Image generation is not deterministic, so a cache hit returns an acceptable image for that input, not a bit-exact regeneration. For most products that is the point. Store outputs in your own storage immediately; do not depend on a returned URL staying valid, and serve from a CDN you control.
Lever 5: spend caps and guardrails
- Set
monthly_spend_cap_usdon every key. Reaching it returns402withspend_cap_exceeded, so a runaway loop becomes an error instead of a bill. - Use a
model_allowliston production keys so premium models cannot be selected by accident or by user-supplied input. - Use separate keys for development and production, with a tight cap on development.
- If end users trigger generations, enforce your own per-user quota before calling the API. Pass the
userfield (up to 64 characters) so usage can be attributed per end user.
Lever 6: pick the host, not just the model
The same model can be served by more than one host at different prices. An unpinned model id routes to the cheapest healthy host, which is usually what you want for cost; pin a host only when you need one for a reason other than price.
Common ways spend leaks
- Silent upscaling: a default resolution set high in a shared helper, applied to thumbnails.
- Retry storms: a client that retries on every error, including ones that will never succeed.
- Unbounded
n: a UI that asks for many variants per click. - Regenerating on page load: an image requested on render instead of generated once and stored.
Each of these is cheap to find if you log model, size and usage.cost on every call, and expensive to find if you do not.
A weekly cost review
From the usage.cost you logged, compute: spend per model, cost per accepted image (including rejected drafts), cache hit rate, and share of spend at the final tier. If final-tier spend is high relative to accepted images, your draft stage is not filtering enough. Choosing the models for each tier is covered in the model-family guide; the pricing page has the current numbers and the quickstart the base request. Create a key and set a cap before you send your first batch.
Frequently asked questions
How is image generation billed?
Per image, with the cost reported in the response's usage.cost field including the platform fee. The pricing basis varies by model: flat per size, per token, per megapixel or per resolution tier.
Are failed image requests charged?
A request that fails upstream is not billed.
What is the easiest way to cut image costs?
Draft on a cheap model at a small size and generate only approved prompts on a premium model at final resolution, and cache outputs so you do not regenerate the same image.
How do I stop a bug from running up a large bill?
Set monthly_spend_cap_usd on each key. When it is reached, requests return 402 spend_cap_exceeded.
Keep reading
- Choosing an Image Generation API — Price, Quality and Control
- Text-to-Image vs Image-to-Image API: Modes, Fields, Pitfalls
- GPT Image vs FLUX vs Seedream vs Nano Banana: How to Choose
- GPT Image API Parameters: Sizes, Quality, Edits and Output
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →