Skip to main content

Overview

MAI-Image 2.6 is Microsoft AI’s in-house image generation model, released on 2026-09-04 and available in public preview on Microsoft Foundry. At launch it ranked No. 2 for both text-to-image and image editing on Arena, and No. 1 for image editing on Artificial Analysis (as of 2026-09-04, per Microsoft’s announcement). APIYI serves two variants through Microsoft’s official channel. Both share the same endpoints and parameters:
  • MAI-Image-2.6: the flagship, tuned for quality and precision
  • MAI-Image-2.6-Flash: the fast variant. Microsoft says it generates 2.8× faster than GPT-Image-2-Medium, and it suits high-throughput production workloads
Highlights: excellent Chinese text rendering (shop signs, vertical couplets, and handwriting come out character-accurate), high-fidelity editing (only the requested part changes; the rest stays pixel-identical), any canvas you want via width + height (up to a 1536×1536 area), and flat per-image pricing regardless of size. A 1024×1024 image takes about 17 s on Flash and about 30 s on 2.6.
📌 Three things to know before you start
  1. Only two endpoints are supported: /v1/images/generations (text-to-image, JSON) and /v1/images/edits (editing, multipart/form-data). /v1/chat/completions and /v1/responses are not supported and return 404.
  2. Do not send response_format, seed, or negative_prompt. All three return 400 immediately. The response is always data[0].b64_json (PNG).
  3. Set the size with width + height, not size. On the text-to-image endpoint size is silently ignored and you always get 1024×1024.
All image APIs are synchronous: there is no async task ID. If the client disconnects, the result is lost but the request is still billed. Give this model a generous timeout. See Image API Essentials & Best Practices.

Text-to-Image API

Generate images from a text prompt, with an interactive Playground.

Image Editing API

Upload a reference image plus an instruction; supports two-image fusion. Includes a Playground.

Let an AI Agent Integrate It for You

If you build with Codex / Claude Code / Cursor, copy the prompt below into it. The agent first fetches the plain-text version of this page (append .md to any docs URL), then writes code for your stack. The common pitfalls are spelled out: timeouts, the three parameters that return 400, width/height instead of size, and file-upload-only editing.

Have a coding agent integrate or debug MAI-Image 2.6 text-to-image and image editing. Copy and paste it into Codex, Claude Code, Cursor, etc.

Why Use MAI-Image 2.6 on APIYI

Official Microsoft Channel

Served through Microsoft’s official channel. The model is the same as the one of the same name on Microsoft Foundry. Standard /v1/images/generations and /v1/images/edits endpoints, with responses shaped like the OpenAI Images API.

Per-Image Pricing

The provider bills by tokens, so larger images cost more. APIYI charges a flat price per image regardless of size: 768×768 and 1536×1536 cost the same, so you can budget per image.

Access From Anywhere

No Azure account or overseas server needed. Reach api.apiyi.com directly from data centers, home networks, or overseas nodes, with one key for every model.

Full Model Lineup

Combine with GPT-Image-2, Nano Banana 2, Seedream, and FLUX for different use cases.

Key Features

Chinese Text Rendering

Chinese shop signs, vertical couplets, and chalkboard handwriting come out character-accurate. Good for posters, product images, and merchandise

High-Fidelity Editing

“Make the teapot cobalt blue” changes only the teapot; dimension labels and other objects stay pixel-identical

Custom Canvas

Any width + height combination, with the long side up to 3072 (e.g. a 3072×768 banner) and an area cap of 1536×1536

Two Speed Tiers

1024×1024 takes about 17 s on Flash and about 30 s on 2.6; latency holds steady at 10 concurrent requests

Sample Results

Chinese text rendering (MAI-Image-2.6-Flash, prompt asked for a Chinese sign welcoming visitors to APIYI): the sign, lanterns, vertical couplets, and chalkboard all show legible Chinese.
MAI-Image-2.6-Flash Chinese text rendering: a traditional teahouse with a Chinese welcome sign
Reference-image editing (MAI-Image-2.6-Flash, instruction “Change the teapot to a deep cobalt blue glaze, keep everything else identical”): original on the left, result on the right. Only the teapot changes color; the dimension labels and other objects are untouched.
MAI-Image-2.6-Flash editing example: teapot recolored from cream to cobalt blue, everything else unchanged

Pricing

Model prices may change; the table above is for reference only, and the Model Pricing tab in the top navigation is authoritative: Model Pricing.
Billing notes
  • Per image, regardless of size: 768×768 and 1536×1536 cost the same, and prompt length does not affect the price.
  • Editing costs the same as text-to-image: single-image edits and two-image fusion are each billed as one image; n=2 on the editing endpoint is billed as 2 images.
  • Requests that fail with 400 (moderation or invalid parameters) produce no image.
  • Do not reconcile with the usage field in the response: prompt_tokens is always 1000 × the image count, a placeholder. The console bill is authoritative.
  • Stacks with the top-up bonus promotion.

Groups and Tokens

This series is in the Default group. Any newly created token can call it; no application is needed.
Token billing mode: both Pay-as-you-go Priority and Per-request work for this series. We recommend Pay-as-you-go Priority, so the same token also works with the token-billed models on the platform.Rate: keep a single key under 50 RPM. For large batch workloads, contact support in advance.

Technical Specs

Endpoints

❌ Chat endpoints are not supported/v1/chat/completions and /v1/responses return 404 Requested path is not found for this series. Chat clients such as Cherry Studio and LobeChat send chat requests to every model in the list, so do not pick MAI-Image in those clients. Use a tool that supports the Images API, or call it from your own code.
✅ The editing endpoint only accepts multipart file uploadsSending JSON (with image as a URL, data URI, or raw base64) to /v1/images/edits returns 400:
Upload the local file directly with -F "[email protected]". No image hosting is needed. See Image Editing API for full examples.
Primary domain https://api.apiyi.com, backup domain https://b.apiyi.com.

Key Parameters

width and height (output size)

Common canvas sizes (all within the area cap):
size behaves differently on the two endpoints: on text-to-image it is silently ignored (always 1024×1024), while on the editing endpoint it does take effect. To avoid confusion, use width + height on both endpoints.

n (image count)

  • Text-to-image: n has no effect. Sending 2, 4, or 10 still returns 1 image (and bills 1). Send parallel requests for more.
  • Editing: n works. n=2 returns 2 images, billed as 2.

Best Practices

1

Pick the variant by use case

Batch generation or latency-sensitive work → MAI-Image-2.6-Flash. Hero posters, complex compositions, or high quality bars → MAI-Image-2.6. Parameters are identical, so switching is just a model-name change.
2

Quote the text you want rendered

Put any text that should appear in the image in quotes and say where it goes, e.g. a sign reading “Grand Opening”. The model reproduces quoted text very faithfully.
3

Say 'keep everything else unchanged' when editing

Write instructions like “Make the teapot cobalt blue, keep everything else exactly the same” to preserve as much of the original as possible.
4

Changing the canvas recomposes the image

If you pass a width / height with a different aspect ratio from the original, the model re-lays out the scene instead of cropping or padding. For local edits, omit the size and the output follows the original’s ratio snapped to multiples of 16 (e.g. a 1344×756 input → 1360×768 output).
5

Need several images? Send parallel requests

Text-to-image returns one image per call, so send 4 parallel requests for 4 images. At 10 concurrent requests, latency matched single requests in our tests.

Error Codes and Retries

Client advice: the 4xx / 500 errors above are deterministic, so retrying is pointless; alert on them instead. Only network timeouts and 429 are worth retrying, with exponential backoff and at most 3 attempts. Keep in mind that requests dropped by a client timeout are still billed, so raise the timeout first.

FAQ

This series only returns b64_json and does not accept the response_format parameter. Even "b64_json" returns 400 Invalid parameters: response_format.Code migrated from gpt-image / DALL·E often sets it explicitly. Remove it; the image is still in data[0].b64_json. The same applies to seed and negative_prompt.
The text-to-image endpoint does not read size. It silently ignores it and renders the default 1024×1024. Use "width": 1536, "height": 1024 instead.On the editing endpoint size does work, but use width + height on both for consistency.
No. The editing endpoint only accepts multipart/form-data file uploads. Passing a URL, data URI, or base64 string as image returns 400.If you only have a URL, download it on your server first, then upload it:
Name the second image’s field image2:
The OpenAI SDK’s client.images.edit(image=[f1, f2]) sends both files as image[]. This series does not accept repeated file fields and returns 400. Single-image edits with the SDK work fine.
No. A mask field returns 400. For local changes, describe the area in the prompt, e.g. “Only make the teapot blue, keep everything else exactly the same”. In our tests the model follows such constraints closely.
Not recommended. Those chat clients use /v1/chat/completions, which returns 404 for this series. Use a tool that supports the OpenAI Images API, or call it directly with the code samples in these docs.
Text-to-image always returns 1. Whatever n you send, you get and pay for 1 image. Send parallel requests for more.On the editing endpoint n works: n=2 returns 2 images and is billed as 2.
No. usage.prompt_tokens is always 1000 × the image count and output_tokens is always 0; these are placeholders. This series is billed per image, and the APIYI console bill is authoritative.
This series uses Microsoft’s official content safety policy, which is fairly strict: real celebrities, gore, well-known IP characters (e.g. Disney), and nudity are blocked.A block returns 400 content_safety_violation with the specific reason in the message. Prompt-level blocks usually come back within 5–8 s; a few are applied after generation and take about as long as a normal image. Retrying the same prompt won’t help; rephrase it.
No. Call it as a normal synchronous request and wait for the full response.
The most common cause is wrong model-name case. The name must be exactly MAI-Image-2.6 or MAI-Image-2.6-Flash; mai-image-2.6-flash returns 503.