How to Use GPT Image 2 for Text to Image (Specs, Prompts and Credit Costs)
A practical guide to GPT Image 2 on VdoBloom: its 11 aspect ratios, 1K/2K/4K tiers, exact 3/5/8 credit costs, prompting patterns that work, and when a different model is the smarter pick.
GPT Image 2 is the OpenAI-powered text-to-image model on VdoBloom, and it costs 3 credits at 1K, 5 credits at 2K and 8 credits at 4K — the same price whether you generate from a prompt or edit an existing image. It is the model to reach for when the picture has to obey a long, literal instruction: product layouts, readable text in the image, diagram-like compositions and clean marketing frames. It is also the strictest content filter in the catalog, so it is the wrong model for swimwear, dance or fitness work.
What GPT Image 2 is actually best at
GPT Image 2 reads a prompt more like a brief than a mood board. Where a diffusion-first model averages your words into a vibe, GPT Image 2 tries to satisfy each clause: the object on the left, the label spelled correctly, the background a specific colour, the light coming from a named direction. That makes it the most reliable model on VdoBloom for four jobs:
- Text inside the image. Packaging labels, poster headlines, signage and UI mockups come out spelled correctly far more often than with fast turbo models.
- Multi-part compositions. Prompts with three or more placed elements (“a bottle centre, a sliced lime left, steam rising right”) hold their layout.
- Clean commercial frames. White-background product shots, flat-lay arrangements and editorial covers that need to look deliberate rather than dreamy.
- Odd aspect ratios. It supports 11 ratios including ultrawide 21:9 and the 5:4 and 4:5 crops most models skip.
Its weakness is the mirror image of its strength. It is literal, so a vague prompt produces a competent but flat picture, and its content filter is set to STRICT in the VdoBloom picker. Anything suggestive is refused at the provider, not by us.
Real specs and exact credit costs
These numbers come from the live model config and the credit table, not from a press release. One credit is $0.05 on the Lite plan, which is $15 for 300 credits.
| Setting | What GPT Image 2 supports | Credits per image | Cost at $0.05/credit |
|---|---|---|---|
| 1K resolution | All 11 aspect ratios, including auto | 3 | $0.15 |
| 2K resolution | All 11 aspect ratios except auto | 5 | $0.25 |
| 4K resolution | Every ratio except 1:1 and auto | 8 | $0.40 |
| Image-to-image (1K / 2K / 4K) | Up to 10 reference images | 3 / 5 / 8 | $0.15 / $0.25 / $0.40 |
| Aspect ratios | auto, 1:1, 5:4, 9:16, 21:9, 16:9, 4:3, 3:2, 4:5, 3:4, 2:3 | — | — |
| Content rating | STRICT filter, 4.6 star community rating | — | — |
Two constraints catch people out, and both are provider rules rather than VdoBloom ones. First, choosing the auto aspect ratio forces the output to 1K, so if you pick auto and 4K you pay for the 4K tier and get a 1K-class file — always name a ratio when you want 4K. Second, a square 1:1 image cannot be rendered at 4K; pick 5:4 or 4:5 and crop if you need a big square.
How to prompt GPT Image 2 well
The prompt field accepts a very long brief, so use it. The pattern that works is subject, then placement, then surface and material, then light, then camera, then what to exclude.
- Name the material, not the adjective. “Brushed aluminium with a matte anodised edge” beats “premium looking” every time.
- Put text in quotation marks and say where it sits. The model spells short strings reliably; long paragraphs of in-image text still drift.
- Give one light source a direction. “Single softbox from camera left, soft falloff, no fill” produces a consistent look across a set of images.
- Describe the frame. 35mm, 85mm, top-down flat lay, eye-level three-quarter — these change the output more than any style word.
- Say what you do not want. Literal models respond well to explicit exclusions such as “no props, no reflections on the backdrop, no drop shadow”.
- Iterate at 1K, finish at 4K. Four 1K test shots cost 12 credits; landing the composition first and paying 8 credits once is the cheapest route to a usable file.
A worked example
Say you need a hero image for a cold brew launch, 16:9, with the product name legible.
Prompt: Studio product photograph of a matte black cold brew can standing centre frame on a wet slate slab. The can label reads “MIDNIGHT ROAST” in condensed white type across the middle third. Two ice cubes at the base, one coffee bean left of the can. Single large softbox from camera left, deep shadow to the right, no fill light. Dark charcoal seamless background, shallow depth of field, 85mm lens look. No hands, no straws, no text other than the label.
Run that at 1K and 16:9 for 3 credits. When the layout is right, rerun the identical prompt at 4K, 16:9 for 8 credits. Total spend to a finished hero: around 17 credits if you take three 1K passes first, which is about $0.85. Open the text-to-image tool, pick GPT Image 2 in the model list, and the resolution and ratio controls appear under the prompt box.
When to pick a different model
GPT Image 2 is not the default for everything, and paying 8 credits for a job another model does better is just waste.
- Swimwear, dance, fitness or fashion-forward shoots. The STRICT filter will refuse them. Use a flexible-tier model such as WAN 2.7 Image at 4 credits instead.
- High-volume drafting. If you need forty thumbnails, a 2-credit turbo model like Z-Image Turbo costs a third as much per image.
- Editing photos of real people. An identity-preserving editor holds a face better across edits than a fresh generation does; see the image editor and pick a model tuned for reference images.
- Newer OpenAI tiers. GPT Image 2.5 Flare costs exactly the same 3 / 5 / 8 credits, runs faster, accepts up to 16 reference images and adds ratios such as 27:16 and 9:8. Its Sunburst sibling is the same price again and trades latency for precision. If your shape is on that list, there is no price reason to stay on GPT Image 2.
If you are editing rather than generating, the matching GPT Image 2 image-to-image model takes up to 10 reference images at the same per-image price. New accounts get 10 free credits with no card, which is three 1K generations to test the model before you decide. Full plan maths lives on the pricing page; annual billing drops Lite to $10 a month.
Frequently asked questions
How much does one GPT Image 2 image cost?
Three credits at 1K, five at 2K and eight at 4K. At the Lite rate of $0.05 per credit that is $0.15, $0.25 and $0.40 per image. Image-to-image costs the same as text-to-image at every tier.
Why did my 4K request come back smaller?
You almost certainly used the auto aspect ratio, which forces a 1K output at the provider. Choose an explicit ratio such as 16:9 or 3:2. Note also that 1:1 has no 4K option at all.
Can GPT Image 2 render text in the image?
Yes, and it is one of the better models on VdoBloom for it. Keep strings short, wrap them in quotation marks and say where on the frame they sit. Long paragraphs still degrade.
Does GPT Image 2 allow swimwear or suggestive prompts?
No. It carries the STRICT content label because OpenAI filters those prompts upstream. VdoBloom lists flexible-tier models for swimwear, dance and fitness work; explicit and illegal content stays blocked on every model.
Do my images have a watermark?
Free-tier downloads are watermarked. Any paid plan downloads watermark-free with commercial rights, and one-time credit packs start at $2.49 for 75 credits that never expire.
Is GPT Image 2.5 worth switching to?
For most jobs, yes — the credit price is identical at all three resolutions. Stay on GPT Image 2 only if you need one of its exclusive ratios such as 5:4, 4:5 or 2:3 at 4K.
Ready to try it?
Create your first AI video in minutes — no credit card required.
Start Creating Free →