Home › Guides › Text vs image input

When to send a prompt alone and when to send an image too

Updated 2026-10-02

On this API there is no separate "generate" and "edit" endpoint. Both go to POST /v1/images; the difference is whether the request carries input_references. That makes the switch easy, but it also means a request can silently behave differently from what you meant. This page explains the two modes, the editing workflows that justify the second one, and the places they go wrong.

Text-to-image

A prompt and a model, plus optional size controls:

{
  "model": "openai/gpt-image-1/openai",
  "prompt": "a red panda astronaut floating in space",
  "aspect_ratio": "16:9",
  "resolution": "2K",
  "n": 1
}

Use it when nothing existing must be preserved: concept art, marketing illustrations, placeholders, backgrounds. Size controls are aspect_ratio and resolution, or an explicit size of the form "WxH", which overrides both. Not every model honours them: some bill and generate at a fixed size regardless of what you ask, and an unsupported combination falls back to the model default instead of erroring. Check the returned image dimensions.

Image-to-image

Add one or more reference images. Each is an image_url entry holding either a public URL or a base64 data: URI; there is no multipart upload.

{
  "model": "gemini-2.5-flash-image",
  "prompt": "make this scene look like a watercolor painting",
  "input_references": [
    {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
  ]
}

The prompt now describes a transformation, not a scene. "Change the jacket to dark green, keep everything else" works better than re-describing the whole picture, because the image already carries the content.

Which models accept references

input_references is supported only on edit-capable models and the chat-native image models (the Gemini image models are examples). Others will not take it. VideoRouter's docs also note that one catalogue model, flux-3-image, supports editing with up to 10 reference images. Because support varies by model, verify the capability on the model page before building a flow, and test with a throwaway request. See the models index for what is in the catalogue.

Editing workflows that earn the second mode

OpenAI-specific controls

Some parameters only affect OpenAI models and are ignored elsewhere: quality, output_format, background, output_compression, and moderation (generation only, not edits). For edits on OpenAI models, input_fidelity ("low" or "high") controls how much fine detail from the input image is preserved. Sending these to another model is harmless, but do not expect them to do anything.

Reading the response

{
  "created": 1731430000,
  "data": [{"b64_json": "...", "url": null, "media_type": "image/png"}],
  "usage": {"cost": 0.063}
}

Some models return base64 in b64_json; a few return a real url. Handle both and fetch URL results promptly. The usage.cost field reports what the call cost including the platform fee, so you can reconcile per request. Streaming is not supported: stream: true is rejected with a 400.

Prompting edits well

State what must stay the same as explicitly as what must change: "keep the face, pose and lighting; replace the background with a plain studio backdrop". When you pass several references, say what each is for, such as "the first image is the product, the second is the style to match". Make one change per request where you can. Compound instructions are more likely to produce a partial result, and because each attempt is billed, several narrow edits that you verify one at a time often cost less than one ambitious edit you have to redo.

Pitfalls

SymptomCauseFix
400 when adding referencesModel has no image inputPick an edit-capable model
Output ignores the inputPrompt re-describes a new sceneWrite the change, not the scene
Wrong output sizeModel ignores size fieldsResize after generation, or choose a model that honours them
Request too largeBig base64 images inlineHost the image and pass a URL, or use POST /v1/uploads (50 MB limit, link expires in 30 minutes)
Unexpected spendChained edits each billedCap iterations; see cost control

Attribution and moderation notes

The optional user field (up to 64 characters) records your own stable end-user id, so spend can be attributed per end user, and it is forwarded to OpenAI models for their abuse monitoring. If you accept uploaded images from the public, validate type and size on your side first, and remember that references you pass in are sent to the model host, so check that this is acceptable for the content your users upload.

Choosing

If the output has to match something you already own, use image-to-image. If you are exploring, start with text-to-image on a cheaper model and graduate approved frames into edits. The live table below shows the price spread across catalogue models, and the model-family guide covers picking among them.

ModelCheapest hostPriciest hostCheapest isHosts
black-forest-labs/flux.2-devMachGen
$0.0031 / image
Fal
$0.012 / image
74% lower2
google/nano-banana-2MachGen
$0.034 / image
Fal
$0.08 / image
57% lower2
google/nano-banana-proMachGen
$0.0672 / image
Fal
$0.15 / image
55% lower2
black-forest-labs/flux.2-proDeepInfra
$0.015 / image
Black Forest Labs
$0.03 / image
50% lower3
black-forest-labs/FLUX.1-devDeepInfra
$0.009 / image
SiliconFlow
$0.014 / image
36% lower2
alibaba/wan-2.6Atlas Cloud
$0.021 / image
DeepInfra
$0.03 / image
30% lower2
black-forest-labs/flux.2-maxBlack Forest Labs
$0.07 / image
DeepInfra
$0.1 / image
30% lower3
seedream-5.0-proAtlas Cloud
$0.036 / image
WaveSpeedAI
$0.045 / image
20% lower3
recraft-4.1WaveSpeedAI
$0.04 / image
Pika
$0.042 / image
5% lower2
bytedance/seedream-4.0WaveSpeedAI
$0.027 / image
DeepInfra
$0.04 / image
33% lower4
bytedance/seedream-5.0-liteAtlas Cloud
$0.0315 / image
WaveSpeedAI
$0.035 / image
10% lower4
qwen-image-maxWaveSpeedAI
$0.07 / image
DeepInfra
$0.075 / image
7% lower2

Per image, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.

A working request is in the quickstart; create a key to try both modes on one of your own images.

Frequently asked questions

How do I do image-to-image with this API?

Send the same POST /v1/images request with an input_references array of image_url entries (URL or base64 data URI) and a prompt describing the change.

Do all image models support input_references?

No. Only edit-capable models and chat-native image models do. Check the model page, since an unsupported model rejects the request.

Is there a separate edit endpoint?

No. Generation and editing share /v1/images; the presence of input_references selects image-to-image.

Why is my output the wrong size?

Some models ignore aspect_ratio, resolution and size and generate at a fixed size. Check the returned dimensions and resize afterwards if needed.

Keep reading

Using image generation is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →