Skip to main content
Qwen3.8-Max (qwen3.8-max) is Alibaba Qwen’s new flagship, released August 3, 2026. It is a sparse MoE model with 2.4 trillion total parameters, a 1M context window, 131K max output, and native support for text, image, and video input. APIYI listed it the day it shipped and ran 586 live test calls against it — the capability matrix, parameter behavior, and billing notes on this page all come from those tests rather than from restating official documentation.
Qwen3.8-Max is live on APIYI: model name qwen3.8-max. Thinking is on by default (at the xhigh tier, and thinking tokens are billed as output), so set reasoning_effort="none" explicitly for everyday chat — in testing this took output from roughly 158 tokens down to 5. For the previous generation, see Qwen3.6 series (legacy).

Why this model

17.5% below official

$1.65 input, $4.95 output per 1M tokens versus Alibaba Cloud’s $2/$6. Top-up promotions stack on top.

1M context, verified

Across 8K / 32K / 128K bodies with markers buried mid-document and at the tail, both endpoints recalled all 6/6 exactly. A 128K call takes about 80 seconds.

Three modalities, one model

Text, image, and video input all verified working — no switching between a “long-context model” and a “vision model.”

Much stronger agentic work

FrontierSWE rose from the previous generation’s 40.7 to 73.5, DeepSWE from 21.6 to 56.6. The tool-calling chain is complete, with two-round round-trips verified.

Endpoint support

Pricing

Per 1M tokens, pre-discount list price: Top-up promotions stack on top for a lower effective cost.

Specifications

Official benchmarks: GPQA Diamond 92.6, PaperBench 93.0, OmniDocBench 1.5 92.1, Terminal-Bench 2.1 86.6, OSWorld-Verified 86.1, IFBench 82.8, FrontierSWE 73.5, SWE-bench Pro 67.7.

Controlling thinking (the most important section)

Qwen3.8-Max thinks by default, at the xhigh tier. Thinking tokens are billed as output and frequently account for over 90% of it.

Seven values, four real tiers

The parameter accepts 7 values but maps to only 4 real tiers in practice: Passing max does not think harder than xhigh. Any other value returns a 400 listing the legal set.

How to turn thinking off

enable_thinking: false in extra_body and chat_template_kwargs: {"enable_thinking": false} are equivalent and also work.
max_tokens does not bound thinking tokens. We set max_tokens=1 and were still billed 1,054 output tokens, 1,045 of them thinking.max_tokens only truncates the visible answer. Use reasoning_effort to control cost — do not rely on max_tokens.

thinking_budget has no effect

Passing 128 / 512 / 4096 all behave identically to the low tier; the number itself is ignored. Use reasoning_effort instead.

Code examples

Python (OpenAI SDK compatible)

Image input

Remote image URLs also work on this endpoint — just set url to an https://... address.

Video input

Video understanding took 144–285 seconds per call in testing. Set your client timeout above 300 seconds and prefer streaming or an async task queue.
There is also a frame-sequence form, {"type": "video", "video": [frame1, frame2, ...]}, which requires 4–8000 frames — fewer than 4 returns a 400.

cURL

Tool calling

Tool calling on the Chat Completions endpoint is fully working: single tool, parallel tools, two-round round-trip, picking 1 out of 20 tools, streaming deltas, and parallel_tool_calls: false all verified.
Forced tool calls require thinking off. When tool_choice is "required" or names a specific function, you must also set reasoning_effort="none" — otherwise you get a 400 (tool_choice does not support being set to required or object in thinking mode) or the call is silently skipped.tool_choice set to "auto" / "none" is unaffected. The same applies to n > 1.

Structured output

response_format with json_schema held strictly in testing: nested objects, enums, arrays, and additionalProperties: false all took effect, with no extra fields and no Markdown fences.
Disable thinking for structured output. Same schema, measured side by side:Conformance was identical; cost and latency differ by an order of magnitude.

Context caching

  • Hit threshold around 1,024 tokens: an 818-token prefix missed; 1,070 tokens and up hit
  • Real multi-turn conversations do hit: appending messages turn by turn hit on every round
  • Long documents benefit most: 98.6% cached input at 128K, 99.3% at 32K
Do not use the cache fields in the API response to judge whether a hit occurred. On some routes cache_read_input_tokens is always 0, and on others the response carries no cache fields at all — yet the very same requests show real cache reads in the console billing records.Trust the “cache billing detail” in the console, which lists token counts and amounts separately for cache creation (1.25x) and cache reads (0.125x).
Cache billing semantics are not fixed per endpoint — they vary by route. In testing, identically shaped requests on the same route were billed under two different schemes on different days:Open a single request in the console log and the “cache billing detail” states which scheme applied and shows the full calculation. That is the only place to find out how a given call was actually billed.
On the Anthropic endpoint, an implicit cache hit is possible without cache_control. Before adding the marker everywhere, compare actual charges in the console — do not assume marking it is always cheaper.

Using the Anthropic endpoint

/v1/messages works for code integration, but you must strip thinking blocks before replaying history, otherwise you get a 400 (if content is list. item must be dict and key[type] should in dict).
With that filter in place, we verified 3-turn cross-turn memory, a two-round tool round-trip, and tool results persisting into later turns.

2026-08-08 follow-up: multi-turn tool calling itself is fine

A dedicated 300-call retest confirms the multi-turn tool_use / tool_result chain itself works. The only blocker is the thinking block:
  • tool_result has no extra format restrictions. String or block-array content, is_error true or false, empty results, 50KB results, out-of-order replay, partial replay, fabricated tool_use_id — all 15 shapes passed. Control characters, emoji, and a 200,000-character single line also passed.
  • No signature value saves you. Empty string, null, the key removed entirely, or a fabricated value all return the same 400. You have to drop the whole block.
  • Stress test passes once stripped. An autonomous agent loop with a 24K-token system prompt and 8 tools, 12 turns × 2 runs, context growing to 28.7K — 24/24 succeeded.
  • SSE events are complete: message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop, plus ping. text_delta, thinking_delta, signature_delta, and input_json_delta all behave correctly.
  • No rate or concurrency limit observed: the same request repeated 40 times sequentially all succeeded, and concurrency levels of 1 / 4 / 8 / 16 / 32 all succeeded with no 429s.
Off-the-shelf clients such as Claude Code are not usable yet. They replay history content blocks verbatim and their behavior cannot be changed, so the first turn returns tool_use normally, then the second turn returns 400 once you send tool_result back — this is the most common failure report on this endpoint.Newer Claude Code builds also send thinking: {"type": "adaptive"}; some routes accept only enabled / disabled / auto and will return 400 on the first turn.Use /v1/chat/completions instead.

What to do if you want it inside Claude Code

For this class of “doesn’t work in one specific client” problem, the limitation may well be on the model side rather than in our adaptation. We suggest verifying the same usage on Alibaba Cloud’s own Bailian platform first (console: bailian.console.aliyun.com):
  • if the official platform rejects it too, it is a model-side limitation and there is nothing we can route around;
  • if it works there but not here, send us the request body and we will take it up with the channel provider.
If your goal is simply to get work done inside Claude Code and similar clients, the APIYI Claude series or OpenAI series is the easier path — the default group is officially routed and needs no extra adaptation.

Other differences and field notes

  • response_format is silently ignored (force a tool call for structured output)
  • tool_choice only accepts the OpenAI format; forced tool calling (required or a named function) is unsupported in thinking mode on both endpoints
  • Images must be base64; remote URLs return 400
  • reasoning_effort has no effect — use thinking: {"type": "disabled"} to turn thinking off
  • stop_sequences does truncate, but stop_reason is misreported as end_turn and the stop_sequence field comes back null, so do not rely on it to detect why generation stopped
  • Streaming usage varies by route: on some routes the input_tokens in message_start is unreliable, and on others the final streamed output_tokens is always 0. For exact accounting, use the non-streaming usage or your billing records
  • Measured input ceiling is 983,616 tokens; going over returns Range of input length should be [1, 983616]
Set generous timeouts. The first SSE byte took 6–17 seconds in testing, with the connection completely silent until then, and larger request bodies are slower still — roughly 44 seconds at 256KB and 160 seconds at 1MB. Behind Docker, a bastion host, or a corporate gateway, an idle timeout at any hop shows up as “hangs for a long time, then exits with an error.” Set the client timeout to 300 seconds or more.

Parameter compatibility

Best practices

Everyday chat and high-volume calls

Set reasoning_effort="none" explicitly. Measured latency dropped from ~5 s to 2 s, and output tokens to roughly 1/30.

Long documents and codebases

128K recall was exact in testing, and long-document cache hit rates are high. Put the large document early in the message list and the question at the tail.

Data extraction

Constrain with json_schema and disable thinking. Conformance is unaffected.

Agents and tool orchestration

Use /v1/chat/completions. Remember to disable thinking when forcing a tool call.

FAQ

max_tokens bounds only the visible answer, not the thinking portion. We measured 1,054 output tokens billed at max_tokens=1. Use reasoning_effort="none" to control cost.
Forced tool choice is not supported while thinking is on. Pass reasoning_effort="none" alongside it.
This endpoint is not wired up for the model yet — all 30 test calls failed, with the error code alternating between 404 and 400. It has been reported upstream, and we will announce it in Live Updates once it is available. Use /v1/chat/completions instead.
Not yet. The /v1/messages endpoint rejects history messages containing thinking blocks, and Claude Code replays them verbatim — so the first turn produces tool_use and the second turn returns 400 once tool_result goes back. When calling from your own code, strip those blocks and the endpoint works fine.If you need to get work done inside Claude Code, the APIYI Claude series or OpenAI series is the easier path — the default group is officially routed and needs no extra adaptation. You can also verify the same usage on Alibaba Cloud’s Bailian platform (bailian.console.aliyun.com) first; if the official platform rejects it too, it is a model-side limitation.
This is the classic symptom on /v1/messages. The replayed assistant message carries a thinking block, which the endpoint rejects with a 400. Setting signature to an empty string or null, or removing the field, does not help — you must drop the entire thinking block.Once stripped, a 12-turn tool loop at 24K context ran to completion in testing. The multi-turn tool_use / tool_result chain itself is not the problem.
The model is served over more than one upstream route and they do not report the same usage fields: some omit reasoning_tokens and cached_tokens, some always report cache_read_input_tokens as 0, and some always report a final streamed output_tokens of 0. This has been reported upstream for alignment.What the API reports is not what you are billed. For exact accounting, use the billing detail on the individual request in the console log, which shows the full calculation for both base and cache charges.
Video understanding measured 144–285 seconds per call; this is the model’s own processing time. Set your timeout above 300 seconds and consider an async queue.
Measurements on this page come from 586 live calls on 2026-08-03 (12:50–14:35 UTC+8). Billing-related conclusions are based on the usage fields returned by the API and were not cross-checked line by line against invoices. Model and gateway behavior may change as channels are adjusted — treat live calls as the source of truth.