When to send a prompt alone and when to send an image too
Updated 2026-10-02
On this API there is no separate "generate" and "edit" endpoint. Both go to POST /v1/images; the difference is whether the request carries input_references. That makes the switch easy, but it also means a request can silently behave differently from what you meant. This page explains the two modes, the editing workflows that justify the second one, and the places they go wrong.
Text-to-image
A prompt and a model, plus optional size controls:
{
"model": "openai/gpt-image-1/openai",
"prompt": "a red panda astronaut floating in space",
"aspect_ratio": "16:9",
"resolution": "2K",
"n": 1
}
Use it when nothing existing must be preserved: concept art, marketing illustrations, placeholders, backgrounds. Size controls are aspect_ratio and resolution, or an explicit size of the form "WxH", which overrides both. Not every model honours them: some bill and generate at a fixed size regardless of what you ask, and an unsupported combination falls back to the model default instead of erroring. Check the returned image dimensions.
Image-to-image
Add one or more reference images. Each is an image_url entry holding either a public URL or a base64 data: URI; there is no multipart upload.
{
"model": "gemini-2.5-flash-image",
"prompt": "make this scene look like a watercolor painting",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
]
}
The prompt now describes a transformation, not a scene. "Change the jacket to dark green, keep everything else" works better than re-describing the whole picture, because the image already carries the content.
Which models accept references
input_references is supported only on edit-capable models and the chat-native image models (the Gemini image models are examples). Others will not take it. VideoRouter's docs also note that one catalogue model, flux-3-image, supports editing with up to 10 reference images. Because support varies by model, verify the capability on the model page before building a flow, and test with a throwaway request. See the models index for what is in the catalogue.
Editing workflows that earn the second mode
- Restyle: keep composition, change look (photo to illustration).
- Local edits: swap colours, objects or backgrounds while keeping the subject.
- Consistency across assets: pass the same character or product image into each request so a set of images shares an identity.
- Variants: generate several crops or colourways from one approved base (
nreturns multiple images per call). - Chained refinement: take the best output, pass it back as the reference with a narrower instruction. Each step is a new, separately billed request.
OpenAI-specific controls
Some parameters only affect OpenAI models and are ignored elsewhere: quality, output_format, background, output_compression, and moderation (generation only, not edits). For edits on OpenAI models, input_fidelity ("low" or "high") controls how much fine detail from the input image is preserved. Sending these to another model is harmless, but do not expect them to do anything.
Reading the response
{
"created": 1731430000,
"data": [{"b64_json": "...", "url": null, "media_type": "image/png"}],
"usage": {"cost": 0.063}
}
Some models return base64 in b64_json; a few return a real url. Handle both and fetch URL results promptly. The usage.cost field reports what the call cost including the platform fee, so you can reconcile per request. Streaming is not supported: stream: true is rejected with a 400.
Prompting edits well
State what must stay the same as explicitly as what must change: "keep the face, pose and lighting; replace the background with a plain studio backdrop". When you pass several references, say what each is for, such as "the first image is the product, the second is the style to match". Make one change per request where you can. Compound instructions are more likely to produce a partial result, and because each attempt is billed, several narrow edits that you verify one at a time often cost less than one ambitious edit you have to redo.
Pitfalls
| Symptom | Cause | Fix |
|---|---|---|
| 400 when adding references | Model has no image input | Pick an edit-capable model |
| Output ignores the input | Prompt re-describes a new scene | Write the change, not the scene |
| Wrong output size | Model ignores size fields | Resize after generation, or choose a model that honours them |
| Request too large | Big base64 images inline | Host the image and pass a URL, or use POST /v1/uploads (50 MB limit, link expires in 30 minutes) |
| Unexpected spend | Chained edits each billed | Cap iterations; see cost control |
Attribution and moderation notes
The optional user field (up to 64 characters) records your own stable end-user id, so spend can be attributed per end user, and it is forwarded to OpenAI models for their abuse monitoring. If you accept uploaded images from the public, validate type and size on your side first, and remember that references you pass in are sent to the model host, so check that this is acceptable for the content your users upload.
Choosing
If the output has to match something you already own, use image-to-image. If you are exploring, start with text-to-image on a cheaper model and graduate approved frames into edits. The live table below shows the price spread across catalogue models, and the model-family guide covers picking among them.
| Model | Cheapest host | Priciest host | Cheapest is | Hosts |
|---|---|---|---|---|
| black-forest-labs/flux.2-dev | MachGen $0.0031 / image | Fal $0.012 / image | 74% lower | 2 |
| google/nano-banana-2 | MachGen $0.034 / image | Fal $0.08 / image | 57% lower | 2 |
| google/nano-banana-pro | MachGen $0.0672 / image | Fal $0.15 / image | 55% lower | 2 |
| black-forest-labs/flux.2-pro | DeepInfra $0.015 / image | Black Forest Labs $0.03 / image | 50% lower | 3 |
| black-forest-labs/FLUX.1-dev | DeepInfra $0.009 / image | SiliconFlow $0.014 / image | 36% lower | 2 |
| alibaba/wan-2.6 | Atlas Cloud $0.021 / image | DeepInfra $0.03 / image | 30% lower | 2 |
| black-forest-labs/flux.2-max | Black Forest Labs $0.07 / image | DeepInfra $0.1 / image | 30% lower | 3 |
| seedream-5.0-pro | Atlas Cloud $0.036 / image | WaveSpeedAI $0.045 / image | 20% lower | 3 |
| recraft-4.1 | WaveSpeedAI $0.04 / image | Pika $0.042 / image | 5% lower | 2 |
| bytedance/seedream-4.0 | WaveSpeedAI $0.027 / image | DeepInfra $0.04 / image | 33% lower | 4 |
| bytedance/seedream-5.0-lite | Atlas Cloud $0.0315 / image | WaveSpeedAI $0.035 / image | 10% lower | 4 |
| qwen-image-max | WaveSpeedAI $0.07 / image | DeepInfra $0.075 / image | 7% lower | 2 |
Per image, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.
A working request is in the quickstart; create a key to try both modes on one of your own images.
Frequently asked questions
How do I do image-to-image with this API?
Send the same POST /v1/images request with an input_references array of image_url entries (URL or base64 data URI) and a prompt describing the change.
Do all image models support input_references?
No. Only edit-capable models and chat-native image models do. Check the model page, since an unsupported model rejects the request.
Is there a separate edit endpoint?
No. Generation and editing share /v1/images; the presence of input_references selects image-to-image.
Why is my output the wrong size?
Some models ignore aspect_ratio, resolution and size and generate at a fixed size. Check the returned dimensions and resize afterwards if needed.
Keep reading
- Choosing an Image Generation API — Price, Quality and Control
- GPT Image vs FLUX vs Seedream vs Nano Banana: How to Choose
- Image Generation API Cost Control: Billing, Drafts, Caching
- GPT Image API Parameters: Sizes, Quality, Edits and Output
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →