GPT Image 2 Image-to-Image vs Text-to-Image: Which Mode to Use
GPT Image 2 image-to-image and text-to-image cost the same 3, 5 and 8 credits at 1K, 2K and 4K. Real specs, the auto-aspect 1K trap, prompting rules and a worked example.
GPT Image 2 image-to-image and GPT Image 2 text-to-image are the same OpenAI-built model behind two different doors on VdoBloom, and they cost exactly the same: 3 credits at 1K, 5 credits at 2K and 8 credits at 4K per image. The only real difference is what you feed them. Text-to-image starts from a blank canvas and a prompt. Image-to-image starts from up to 10 photos you upload and rewrites them. Because the price is identical at every resolution, the choice is never about money β it is about whether you already own the subject in the frame.
The same engine, two entry points
Inside VdoBloom these are two separate model IDs, gpt-image-2-text-to-image and gpt-image-2-image-to-image, and they live on two different tabs. Text-to-image sits on the image generation tab. Image-to-image sits on the image editing tab, where the upload area appears first and the prompt box second. Everything else β the eleven aspect ratios, the 1K/2K/4K resolution ladder, the 20,000-character prompt limit β is shared.
That shared spec sheet is unusual. Most model families give the editing variant a shorter prompt limit or fewer aspect ratios. GPT Image 2 does not. If you can describe it for a fresh generation, you can describe the same change over a reference photo, in the same number of words.
Spec and price, side by side
| Spec | GPT Image 2 text-to-image | GPT Image 2 image-to-image |
|---|---|---|
| Input | Prompt only | 1β10 reference images plus prompt |
| Max upload size | Not applicable | 30 MB per image |
| Prompt limit | 20,000 characters | 20,000 characters |
| Aspect ratios | auto, 1:1, 5:4, 9:16, 21:9, 16:9, 4:3, 3:2, 4:5, 3:4, 2:3 | Identical list of 11 |
| Resolutions | 1K / 2K / 4K | 1K / 2K / 4K |
| 1K cost | 3 credits ($0.15) | 3 credits ($0.15) |
| 2K cost | 5 credits ($0.25) | 5 credits ($0.25) |
| 4K cost | 8 credits ($0.40) | 8 credits ($0.40) |
| Tab | /dashboard/images/generate/ | /dashboard/images/edit/ |
Dollar figures use the Lite plan rate, where $15 buys 300 credits and one credit is therefore $0.05. On that plan a month of Lite is 100 images at 1K, 60 at 2K or 37 at 4K from either mode. The full ladder, including the annual rate and the one-time credit packs that never expire, is on the pricing page.
Two resolution rules that catch people out
The provider enforces two constraints, and both apply to text-to-image and image-to-image equally. First, choosing the auto aspect ratio forces the output to 1K β auto lets the model decide the frame, and it decides at the base tier. VdoBloom bills an unspecified resolution at the 1K rate of 3 credits to match what actually comes back, so you are not charged 4K money for a 1K file. Second, the square 1:1 ratio cannot be rendered at 4K. If you need a 4K square, generate at 4K in 5:4 or 4:3 and crop, or stay at 2K.
The practical habit: pick a real aspect ratio rather than leaving it on auto whenever resolution matters to you. Auto is fine for quick thumbnails and mood exploration; it is the wrong setting for anything you intend to print or upscale.
How to prompt each one
For text-to-image, the model rewards scene grammar: subject, then what the subject is doing, then lighting, then lens and framing, then style. With a 20,000-character budget you can afford a full paragraph per layer, and GPT Image 2 genuinely uses it. Its strongest single trait is legible text inside the image β signage, packaging copy, poster headlines β so put the exact words you want rendered in quotation marks in the prompt.
For image-to-image the grammar flips. Lead with what must NOT change, then name the single change. Something like: keep the subject, pose, face and packaging text exactly as in the reference; replace only the background with a sunlit marble counter. Long descriptive prompts that redescribe the whole scene push the model toward re-inventing your subject, which is the most common complaint about any editing model. Uploading several references of the same product from different angles makes identity hold up far better than uploading one.
If you are editing a photo that contains a real person, you must have that personβs consent before uploading it.
A worked example
Say you are building a launch post for a coffee brand. Text-to-image at 2K in 4:5, 5 credits, gives you the stylised hero scene β the pour, the steam, the typography on a fictional bag. That is the right mode because nothing in the frame exists yet.
Now you have real product photography of the actual bag. Text-to-image is now the wrong tool: it would invent a bag that is close but not yours. Switch to image-to-image, upload three angles of the real bag, and prompt for the same marble-counter scene with the packaging preserved. At 2K that is another 5 credits. Four background variations cost 20 credits, roughly a dollar, and every one keeps the real label.
When you want the same subject in many poses rather than many backgrounds, the pose variations tab runs GPT Image 2 image-to-image across a batch for you instead of you writing each prompt by hand.
When to pick a different model
GPT Image 2 has two successors on the platform at the same price. GPT Image 2.5 Flare is the fast variant and GPT Image 2.5 Sunburst is the precision variant; both cost the same 3, 5 and 8 credits, and both accept 16 reference images instead of 10. Their aspect ratio list is different, not larger β it drops 5:4, 4:5 and 2:1 and adds 27:16, 16:27, 9:8 and 8:9, several of which are 1K only. If you specifically need 5:4 or 4:5 at 4K, GPT Image 2 is still the one that does it.
Outside the family, choose a cheaper model when the output is disposable: Z-image Turbo at 2 credits and FLUX.2 [dev] at 4 credits are both well under GPT Image 2βs 1K price. Choose a pricier one when the brief is typography-heavy commercial work β Recraft V4 Pro at 22 credits and Riverflow 2.0 Pro at 18 credits exist for that. And if the subject is going to move, generate the still here and take it into a video model; the free AI video generator page explains how the still-to-video handoff works on a new accountβs 10 free credits.
The verdict by use case
Pick text-to-image when the subject does not exist yet, when you want stylistic freedom, or when you are exploring concepts and identity does not matter. Pick image-to-image when you already own the subject β a product, a face, a room, a logo β and the job is to move it somewhere else without it changing. Since both cost the same, the expensive mistake is not picking the wrong mode, it is picking 4K when 1K would have done: that is 8 credits against 3, on every single render.
Frequently asked questions
Does image-to-image cost more than text-to-image on GPT Image 2?
No. Both are priced purely on output resolution: 3 credits at 1K, 5 at 2K, 8 at 4K. Uploading ten reference images costs the same as uploading one.
How many reference images can GPT Image 2 image-to-image take?
Up to 10, each a maximum of 30 MB. GPT Image 2.5 Flare and Sunburst raise that to 16.
Why did my 4K request come back at 1K?
You almost certainly left the aspect ratio on auto, which forces 1K. Pick an explicit ratio. Separately, 1:1 cannot render at 4K at all.
Can I use GPT Image 2 output commercially?
Yes on any paid plan, which also removes the watermark and grants commercial rights. Free accounts can generate with their 10 starter credits but download watermarked.
Is there a prompt length difference between the two modes?
No, both accept up to 20,000 characters. In practice editing prompts should be much shorter than generation prompts, because over-describing the scene makes the model redraw your subject.
Ready to try it?
Create your first AI video in minutes β no credit card required.
Start Creating Free β