Skip to main content

TL;DR

All three models are built on OpenAI’s gpt-image-2 underneath. The differences are in channel nature (official direct vs reverse-engineered), pricing model, and parameter granularity.
The two reverse siblings (-all / -vip): This page’s “Reverse” column covers both gpt-image-2-all and gpt-image-2-vip — they share identical call format and the same $0.03/image flat price. The difference now is speed vs quality:
  • gpt-image-2-all: ChatGPT web line, ~90s generation — speed is the advantage
  • gpt-image-2-vip: Codex line, ~120–200s generation — slower, but sometimes higher quality
  • Both: no quality, no n, no mask inpainting
⚠️ -vip’s size parameter is currently broken (since 2026-06-23, due to a Codex generation-rule change; output is fixed to adaptive 1K, no recovery ETA) — for locked sizes / 4K, use the official gpt-image-2. Same for quality tiers or mask inpainting.
About current speeds: -all / -vip generation is slower than at launch due to OpenAI upstream compute fluctuations — this affects all reverse-channel users, not just APIYI; our account pool and ops are healthy. Set client timeouts to 300s+ and leave more headroom for complex prompts.

Full Comparison Table

🔑 Create or manage API tokens: https://api.apiyi.com/token
When creating a token in the console, choose a group (Default is fine) and a token type (Per-call / Token-priority). Calling gpt-image-2 (official) requires a “Token-priority” token — per-call tokens will be rejected due to billing-mode mismatch.

When to Pick Each

Pick gpt-image-2-all (Reverse) when

💰 Predictable cost

Stable $0.03/image with no size/quality tier. Ideal for batch production with hard cost ceilings (infographics, marketing assets, e-commerce thumbnails).

⚡ Faster output

~90s generation — slightly faster than both -vip and the official version. Better real-time UX.

🔁 One codebase, swap anytime

Standard Images API format — same code as -vip and the official-relay gpt-image-2; switch or fall back by changing the model name.

🌏 Chinese + marketing text

Native Chinese prompt support, excellent text rendering for signage / posters / infographics — great for Chinese-audience content production.

Pick gpt-image-2-vip (Reverse, quality-first) when

🎨 Sometimes higher quality

The Codex line’s detail rendering is occasionally better than -all — for showcase images where you’re not in a hurry and want a bit more quality at the same reverse-channel flat price.

⏱️ Trade time for quality

~120–200s generation, slower than -all — pick it when you can accept a longer wait for a higher ceiling.

🔁 Code shared with -all

Identical request structure to -allone codebase switches between both models by swapping the model name based on your speed / quality preference.

💰 Cost still predictable

Same flat $0.03/image as -all — batch production costs stay capped.
-vip’s former “locked sizes / 4K” selling point does not currently hold: the size parameter has been broken since 2026-06-23 (a Codex generation-rule change; output is fixed to adaptive 1K, no recovery ETA). For e-commerce hero shots, poster templates, 4K wallpapers and any locked-size / 4K needs, use the official gpt-image-2.

Pick gpt-image-2 (Official) when

🎚️ Quality tiers

quality supports low/medium/high/auto. Use low for drafts to save cost; high for print-grade finals — official-only; both reverse models reject it.

🎯 Mask inpainting

Alpha-channel mask supported — precisely modify a region while preserving the rest. Both reverse models do not support this.

🖼️ Locked sizes / 4K

size accepts any valid resolution (including 4K). While the reverse channel’s size is broken, every exact-dimension or 4K workload goes official.

🔌 Same as OpenAI Official

Goes through the official Images API — fields and behavior identical to OpenAI official. Existing OpenAI-SDK-based code / systems migrate with zero changes and stay stable long-term.

Key Differences in Detail

1. b64_json format gotcha (migration trap!)

As verified in July 2026, both models now return raw base64 (no data: prefix) — but gpt-image-2-all used to include the prefix, so the safest shared code checks for it first:
When switching between the two, the b64_json handling code must change, or you’ll get a corrupted data URL or a decode failure.

2. Resolution control

gpt-image-2-all (in the prompt):
gpt-image-2-vip (size currently broken, since 2026-06-23): It used to accept 30 explicit sizes (including 4K), but after a Codex generation-rule change the size parameter is broken and output is fixed to adaptive 1K, with no recovery ETA. For now, describe the composition in the prompt just like -all; for exact dimensions / 4K, use the official gpt-image-2. gpt-image-2 (size parameter strict + quality tiers):

3. Upload / output format differences

4. Cost ballpark

Bottom line: For batch / low-quality workloads, the reverse channel isn’t always cheaper (1K low is actually less expensive on the official tier). The mid-to-high quality range is the reverse channel’s $0.03 sweet spot. Pick official (token-metered) when you need quality tiers / mask inpainting / locked sizes, 4K / strict OpenAI-API field parity.

Client Settings

Common to all three models: for image edit / multi-image fusion, compress each input image to under 1.5MB (JPEG quality 80-90 / down-sized resolution). Sporadic shell_api_error / Unknown error responses are most often triggered by oversized inputs — compressing measurably improves success rate and latency. Output resolution is independent of input size — quality is set on the output side (size + quality for official; prompt phrasing for -all and for -vip while its size is broken), not by input file size.

FAQ

Yes, strongly recommended. For all three models, compress each input image to under 1.5MB (JPEG quality 80-90 / down-sized resolution): sporadic shell_api_error / Unknown error responses are most often triggered by oversized inputs, and compressing measurably improves success rate and latency.Don’t worry about compression hurting quality — output resolution is independent of input size. The “output-side” controls differ across the three:
  • gpt-image-2-all: controlled by prompt composition phrasing (see the verified phrasing table on the -all overview page) — 4K / 8K in the prompt does not count
  • gpt-image-2-vip: the size field is currently broken (fixed to adaptive 1K) — put composition intent in the prompt too
  • gpt-image-2: controlled by size + quality (any valid size)
Bottom line: shrinking inputs only speeds things up — quality is set by output-side configuration, not input file size.
Yes. All three run on the Default channel — the same API Key calls them with no extra config. Note: calling gpt-image-2 (official) requires a “Token-priority” token; -all / -vip accept either token type.
Use the OpenAI Images API (/v1/images/generations for text-to-image + /v1/images/edits for editing), for two reasons:
  1. More stable: upstream resource supply for the Images API channel is more plentiful, so call success rates are higher
  2. Compatible with the official relay for easy switching: the call method and parameter format are fully compatible with the official-relay gpt-image-2 — if the reverse channel hits risk-control turbulence, just swap the model name to switch to the official relay with zero code changes
There is also a chat-based endpoint (/v1/chat/completions, no longer recommended), only useful for multi-turn iterative editing or passing online image URLs directly. Note that when the image intent is ambiguous, it may return plain text instead of an image (prepend a fixed prefix like “Generate an image:” to reinforce it). For full parameters, see the -all chat-based API reference / -vip chat-based API reference.
Both are reverse-engineered channels at the same flat price ($0.03/image), with identical call format (-vip’s size is currently broken, so neither takes size). The difference is speed vs quality:
  • Generation time: -all ~90s — speed is the advantage; -vip ~120–200s. Currently slower than at launch due to OpenAI upstream compute fluctuations
  • Quality: -vip (Codex line) detail rendering is sometimes higher — for showcase images when you’re not in a hurry
Decision: want fast output → -all; quality-first, not in a hurry → -vip; need locked sizes or 4K → official gpt-image-2. See the GPT-Image-2-VIP Overview for details.
Use the official gpt-image-2. -vip’s size parameter has been broken since 2026-06-23 (fixed to adaptive 1K, no recovery ETA), so neither reverse model can precisely control output dimensions right now.Official-only features: any valid size (incl. 4K), quality tiers (low/medium/high/auto), mask inpainting (alpha-channel mask), strict OpenAI-API field parity (zero-change migration for existing OpenAI-SDK code). Token-metered billing.
  • Stick with the OpenAI SDK / must match OpenAI official, or need locked sizes / 4K: pick gpt-image-2 (official). Drop input_fidelity, avoid background: transparent, leave the rest unchanged.
  • Cut cost, want fast output: pick gpt-image-2-all (reverse, ~90s).
  • Cut cost, quality-first and not in a hurry: pick gpt-image-2-vip (reverse, ~120–200s).
Yes. A common pattern: primary -all or -vip (predictable cost — pick by speed / quality preference), fallback gpt-image-2 (switch when you need quality tiers, mask, or locked sizes). The reverse and official response shapes differ — normalize at the business layer.