Skip to main content
POST
Text-to-image: generate an image from a text prompt
The interactive Playground on the right lets you test online. Enter your API key under Authorization (format: Bearer sk-xxx), choose a model, type a prompt, optionally set width / height, and send.
Use case: this page is for generating images from text only. To modify an existing image or fuse two images, use the Image Editing API.
⚠️ Three parameters return 400This series does not accept response_format, seed, or negative_prompt; sending any of them returns 400 Invalid parameters: xxx. If you are migrating from gpt-image / DALL·E, remove response_format first. The response is always data[0].b64_json.
⚠️ Use width + height, not sizeThis endpoint silently ignores size and always renders 1024×1024. Use integer width + height (always together): each side ≥ 768, width × height ≤ 2,359,296 (a 1536×1536 area).
All image APIs are synchronous: there is no async task ID. If the client disconnects, the result is lost but the request is still billed. A 1024×1024 image takes about 17 s on Flash and about 30 s on 2.6. Set the client timeout to at least 120 s for Flash and 180 s for 2.6. See Image API Essentials & Best Practices.

Code Examples

Python (OpenAI SDK)

Python (requests)

cURL

Node.js (fetch)

Need several images: send parallel requests

The text-to-image endpoint returns 1 image per call (n has no effect), so run requests in parallel:

Parameter Reference

quality, output_format, background, style, and other OpenAI-style fields are silently ignored. Output is always PNG.

Response Format

Response fields
  • b64_json is plain base64 without a data:image/png;base64, prefix. Decode it directly to get a PNG.
  • There is no url field, and no revised_prompt.
  • A 1024×1024 PNG is about 1.5–1.7 MB (2.1–2.3 MB as base64); 1536×1536 is about 4–5 MB. Check your client’s response size limits.
Do not reconcile billing with usage: prompt_tokens is always 1000 × the image count and output_tokens is always 0; these are placeholders. This series is billed per image, and the APIYI console bill is authoritative.

Authorizations

Authorization
string
header
required

API key from the APIYI console

Body

application/json
model
enum<string>
default:MAI-Image-2.6-Flash
required

Model ID (case-sensitive). 2.6 favors quality, Flash favors speed

Available options:
MAI-Image-2.6-Flash,
MAI-Image-2.6
prompt
string
required

Prompt in any language. Put text that should appear in the image in quotes

Example:

"A traditional teahouse storefront with a wooden sign that reads \"Welcome\", red lanterns, warm dusk light, photorealistic"

width
integer
default:1024

Output width in pixels. Each side ≥ 768, width × height ≤ 2,359,296 (a 1536×1536 area); must be sent together with height; rounded down to a multiple of 16.

Required range: x >= 768
Example:

1024

height
integer
default:1024

Output height in pixels, same rules as width

Required range: x >= 768
Example:

1024

Response

Image generated

created
integer

Creation timestamp

Example:

1790999642

data
object[]

Image results; always 1 item for text-to-image

usage
object

Placeholder values, not for billing reconciliation. prompt_tokens is always 1000 × images and output_tokens is always 0. The console bill is authoritative.