Skip to main content

TL;DR

All three models are built on OpenAI’s gpt-image-2 underneath. The differences are in channel nature (official direct vs reverse-engineered), pricing model, and parameter granularity.
The two reverse siblings (-all / -vip): This page’s “Reverse” column covers both gpt-image-2-all and gpt-image-2-vip — they share the same call format (-vip additionally supports the size field) and the same $0.03/image flat price. The difference now is speed vs quality + size locking:
  • gpt-image-2-all: ChatGPT web line, ~90s generation — speed is the advantage
  • gpt-image-2-vip: Codex line, ~120–200s generation — slower, but sometimes higher quality, and it supports size locking (30 presets incl. 4K, restored since 2026-07-22)
  • Both: no quality, no n, no mask inpainting
For quality tiers, mask inpainting, or arbitrary custom sizes beyond the 30 presets, use the official gpt-image-2.
About current speeds: -all / -vip generation is slower than at launch due to OpenAI upstream compute fluctuations — this affects all reverse-channel users, not just APIYI; our account pool and ops are healthy. Set client timeouts to 300s+ and leave more headroom for complex prompts.

Full Comparison Table

🔑 Create or manage API tokens: https://api.apiyi.com/token
When creating a token in the console, choose a group (Default is fine) and a token type (Per-call / Token-priority). Calling gpt-image-2 (official) requires a “Token-priority” token — per-call tokens will be rejected due to billing-mode mismatch.

When to Pick Each

Pick gpt-image-2-all (Reverse) when

💰 Predictable cost

Stable $0.03/image with no size/quality tier. Ideal for batch production with hard cost ceilings (infographics, marketing assets, e-commerce thumbnails).

⚡ Faster output

~90s generation — slightly faster than both -vip and the official version. Better real-time UX.

🔁 One codebase, swap anytime

Standard Images API format — same code as -vip and the official-relay gpt-image-2; switch or fall back by changing the model name.

🌏 Chinese + marketing text

Native Chinese prompt support, excellent text rendering for signage / posters / infographics — great for Chinese-audience content production.

Pick gpt-image-2-vip (Reverse, quality-first) when

🎨 Sometimes higher quality

The Codex line’s detail rendering is occasionally better than -all — for showcase images where you’re not in a hurry and want a bit more quality at the same reverse-channel flat price.

⏱️ Trade time for quality

~120–200s generation, slower than -all — pick it when you can accept a longer wait for a higher ceiling.

🖼️ Locked sizes / 4K

The size parameter is restored (since 2026-07-22): 30 preset sizes (10 ratios × 1K/2K/4K). E-commerce hero shots, poster templates and 4K wallpapers come out at exact dimensions — flat $0.03/image, no 4K surcharge.

🔁 Code shared with -all

Same request structure as -all (just one extra size field) — one codebase switches between both models by swapping the model name based on your speed / quality preference.
-vip’s size only works on the /v1/images/generations and /v1/images/edits endpoints — the /v1/chat/completions chat endpoint does not support size. For arbitrary custom sizes beyond the 30 presets, quality tiers, or mask inpainting, use the official gpt-image-2. Availability of this parameter follows upstream changes — see Live Updates for the latest status.

Pick gpt-image-2 (Official) when

🎚️ Quality tiers

quality supports low/medium/high/auto. Use low for drafts to save cost; high for print-grade finals — official-only; both reverse models reject it.

🎯 Mask inpainting

Alpha-channel mask supported — precisely modify a region while preserving the rest. Both reverse models do not support this.

🖼️ Arbitrary custom sizes

size accepts any valid resolution (including 4K), not limited to presets. -vip only supports 30 preset sizes — anything beyond those 30 goes official.

🔌 Same as OpenAI Official

Goes through the official Images API — fields and behavior identical to OpenAI official. Existing OpenAI-SDK-based code / systems migrate with zero changes and stay stable long-term.

Key Differences in Detail

1. b64_json format gotcha (migration trap!)

As verified in July 2026, both models now return raw base64 (no data: prefix) — but gpt-image-2-all used to include the prefix, so the safest shared code checks for it first:
When switching between the two, the b64_json handling code must change, or you’ll get a corrupted data URL or a decode failure.

2. Resolution control

gpt-image-2-all (in the prompt):
gpt-image-2-vip (size restored, since 2026-07-22): Accepts 30 preset sizes (10 ratios × 1K/2K/4K) — pass size: "WIDTHxHEIGHT" directly (must be one of the 30 presets; full list in the 30-size table):
gpt-image-2 (size parameter strict + quality tiers):

3. Upload / output format differences

4. Cost ballpark

Bottom line: For batch / low-quality workloads, the reverse channel isn’t always cheaper (1K low is actually less expensive on the official tier). The mid-to-high quality range is the reverse channel’s $0.03 sweet spot. Pick official (token-metered) when you need quality tiers / mask inpainting / locked sizes, 4K / strict OpenAI-API field parity.

Client Settings

Common to all three models: for image edit / multi-image fusion, compress each input image to under 1.5MB (JPEG quality 80-90 / down-sized resolution). Sporadic shell_api_error / Unknown error responses are most often triggered by oversized inputs — compressing measurably improves success rate and latency. Output resolution is independent of input size — quality is set on the output side (size + quality for official; the size tier for -vip; prompt phrasing for -all), not by input file size.

FAQ

Yes, strongly recommended. For all three models, compress each input image to under 1.5MB (JPEG quality 80-90 / down-sized resolution): sporadic shell_api_error / Unknown error responses are most often triggered by oversized inputs, and compressing measurably improves success rate and latency.Don’t worry about compression hurting quality — output resolution is independent of input size. The “output-side” controls differ across the three:
  • gpt-image-2-all: controlled by prompt composition phrasing (see the verified phrasing table on the -all overview page) — 4K / 8K in the prompt does not count
  • gpt-image-2-vip: controlled by the size field (30 preset sizes incl. 4K, restored since 2026-07-22) — 4K / 8K in the prompt does not count either
  • gpt-image-2: controlled by size + quality (any valid size)
Bottom line: shrinking inputs only speeds things up — quality is set by output-side configuration, not input file size.
Yes. All three run on the Default channel — the same API Key calls them with no extra config. Note: calling gpt-image-2 (official) requires a “Token-priority” token; -all / -vip accept either token type.
Use the OpenAI Images API (/v1/images/generations for text-to-image + /v1/images/edits for editing), for two reasons:
  1. More stable: upstream resource supply for the Images API channel is more plentiful, so call success rates are higher
  2. Compatible with the official relay for easy switching: the call method and parameter format are fully compatible with the official-relay gpt-image-2 — if the reverse channel hits risk-control turbulence, just swap the model name to switch to the official relay with zero code changes
There is also a chat-based endpoint (/v1/chat/completions, no longer recommended), only useful for multi-turn iterative editing or passing online image URLs directly. Note that when the image intent is ambiguous, it may return plain text instead of an image (prepend a fixed prefix like “Generate an image:” to reinforce it). For full parameters, see the -all chat-based API reference / -vip chat-based API reference.
Both are reverse-engineered channels at the same flat price ($0.03/image), with the same call format (-vip additionally supports size locking). The difference is speed vs quality + size locking:
  • Generation time: -all ~90s — speed is the advantage; -vip ~120–200s. Currently slower than at launch due to OpenAI upstream compute fluctuations
  • Quality: -vip (Codex line) detail rendering is sometimes higher — for showcase images when you’re not in a hurry
  • Size locking: -vip supports 30 preset size values (incl. 4K); -all rejects size — composition goes into the prompt
Decision: want fast output → -all; quality-first or need locked sizes / 4K → -vip; need custom sizes beyond the 30 presets, quality tiers, or mask → official gpt-image-2. See the GPT-Image-2-VIP Overview for details.
Start with gpt-image-2-vip: its size parameter was restored on 2026-07-22 and supports 30 preset sizes (10 ratios × 1K/2K/4K) at a flat $0.03/image with no 4K surcharge. Note that size only works on the images endpoints and must be one of the 30 presets.Go official (gpt-image-2, token-metered) when you need any valid size beyond the 30 presets, quality tiers (low/medium/high/auto), mask inpainting (alpha-channel mask), or strict OpenAI-API field parity (zero-change migration for existing OpenAI-SDK code).
  • Stick with the OpenAI SDK / must match OpenAI official, or need custom sizes beyond the 30 presets: pick gpt-image-2 (official). Drop input_fidelity and leave the rest unchanged (background: transparent keeps working).
  • Cut cost, want fast output: pick gpt-image-2-all (reverse, ~90s).
  • Cut cost, quality-first or need locked sizes / 4K: pick gpt-image-2-vip (reverse, ~120–200s, 30 preset sizes incl. 4K).
Yes. A common pattern: primary -all or -vip (predictable cost — pick by speed / quality preference), fallback gpt-image-2 (switch when you need quality tiers, mask, or custom sizes beyond the 30 presets). The reverse and official response shapes differ — normalize at the business layer.