Home › Guides › Cost control

Keeping per-image spend predictable at volume

Updated 2026-10-02

Image generation looks cheap per call and gets expensive through multiplication: retries, variants, regenerations, and a few high-resolution requests hidden in a stream of cheap ones. Cost control is mostly bookkeeping plus a handful of levers. This page covers how the billing works and how to use each lever.

How billing works

ModelCheapest hostPriciest hostCheapest isHosts
black-forest-labs/flux.2-devMachGen
$0.0031 / image
Fal
$0.012 / image
74% lower2
google/nano-banana-2MachGen
$0.034 / image
Fal
$0.08 / image
57% lower2
google/nano-banana-proMachGen
$0.0672 / image
Fal
$0.15 / image
55% lower2
black-forest-labs/flux.2-proDeepInfra
$0.015 / image
Black Forest Labs
$0.03 / image
50% lower3
black-forest-labs/FLUX.1-devDeepInfra
$0.009 / image
SiliconFlow
$0.014 / image
36% lower2
alibaba/wan-2.6Atlas Cloud
$0.021 / image
DeepInfra
$0.03 / image
30% lower2
black-forest-labs/flux.2-maxBlack Forest Labs
$0.07 / image
DeepInfra
$0.1 / image
30% lower3
seedream-5.0-proAtlas Cloud
$0.036 / image
WaveSpeedAI
$0.045 / image
20% lower3
recraft-4.1WaveSpeedAI
$0.04 / image
Pika
$0.042 / image
5% lower2
bytedance/seedream-4.0WaveSpeedAI
$0.027 / image
DeepInfra
$0.04 / image
33% lower4
bytedance/seedream-5.0-liteAtlas Cloud
$0.0315 / image
WaveSpeedAI
$0.035 / image
10% lower4
qwen-image-maxWaveSpeedAI
$0.07 / image
DeepInfra
$0.075 / image
7% lower2

Per image, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.

Lever 1: resolution and quality

Where a model honours size controls, the price often rises with resolution. Generate at the smallest size that satisfies the display target, and request higher resolution only for final assets. For OpenAI models the quality setting is an additional lever, applied only to those models. Check on the model page whether size actually changes the price, because on models that ignore size controls, asking for less saves nothing.

Lever 2: draft, then final

Run exploration on a cheap, fast model and promote only approved prompts to a premium model. Because the endpoint is shared, promotion changes only the model string. A simple workflow:

  1. Generate 4 draft variants on the draft model.
  2. A person or automated check picks one.
  3. Re-run that prompt (and seed-like inputs) once on the final model at final resolution.

This turns the expensive model's cost into a function of accepted images rather than attempted ones. For edits, use the approved draft as the reference image on an edit-capable model; see image modes.

Lever 3: batching and concurrency

For bulk jobs, queue the work and cap concurrency per model. This does not change the unit price, but it bounds how fast you can spend and keeps you under per-key rate limits (a 429 includes Retry-After). Retry only transient failures (429, 5xx) with backoff, and never retry a 400, which will fail identically every time. If your volume is large and latency-insensitive, check the platform's Batch documentation for whether a discounted route covers your model, since discounts are not available on every model.

Lever 4: cache outputs

The cheapest image is the one you do not regenerate. Cache by a hash of everything that determines the result:

import hashlib, json

def cache_key(model, prompt, size, refs):
    blob = json.dumps({"m": model, "p": prompt, "s": size, "r": sorted(refs)}, sort_keys=True)
    return hashlib.sha256(blob.encode()).hexdigest()

def get_image(db, store, payload):
    k = cache_key(payload["model"], payload["prompt"],
                  payload.get("size") or (payload.get("aspect_ratio"), payload.get("resolution")),
                  [r["image_url"]["url"] for r in payload.get("input_references", [])])
    if hit := db.get(k):
        return store.url(hit)
    resp = call_images_api(payload)               # your POST /v1/images wrapper
    key = store.save(resp)                        # copy to your own storage
    db.set(k, key)
    return store.url(key)

Image generation is not deterministic, so a cache hit returns an acceptable image for that input, not a bit-exact regeneration. For most products that is the point. Store outputs in your own storage immediately; do not depend on a returned URL staying valid, and serve from a CDN you control.

Lever 5: spend caps and guardrails

Lever 6: pick the host, not just the model

The same model can be served by more than one host at different prices. An unpinned model id routes to the cheapest healthy host, which is usually what you want for cost; pin a host only when you need one for a reason other than price.

Common ways spend leaks

Each of these is cheap to find if you log model, size and usage.cost on every call, and expensive to find if you do not.

A weekly cost review

From the usage.cost you logged, compute: spend per model, cost per accepted image (including rejected drafts), cache hit rate, and share of spend at the final tier. If final-tier spend is high relative to accepted images, your draft stage is not filtering enough. Choosing the models for each tier is covered in the model-family guide; the pricing page has the current numbers and the quickstart the base request. Create a key and set a cap before you send your first batch.

Frequently asked questions

How is image generation billed?

Per image, with the cost reported in the response's usage.cost field including the platform fee. The pricing basis varies by model: flat per size, per token, per megapixel or per resolution tier.

Are failed image requests charged?

A request that fails upstream is not billed.

What is the easiest way to cut image costs?

Draft on a cheap model at a small size and generate only approved prompts on a premium model at final resolution, and cache outputs so you do not regenerate the same image.

How do I stop a bug from running up a large bill?

Set monthly_spend_cap_usd on each key. When it is reached, requests return 402 spend_cap_exceeded.

Keep reading

Using image generation is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →