qwen3.8-max) is Alibaba Qwen’s new flagship, released August 3, 2026. It is a sparse MoE model with 2.4 trillion total parameters, a 1M context window, 131K max output, and native support for text, image, and video input. APIYI listed it the day it shipped and ran 586 live test calls against it — the capability matrix, parameter behavior, and billing notes on this page all come from those tests rather than from restating official documentation.
qwen3.8-max. Thinking is on by default (at the xhigh tier, and thinking tokens are billed as output), so set reasoning_effort="none" explicitly for everyday chat — in testing this took output from roughly 158 tokens down to 5. For the previous generation, see Qwen3.6 series (legacy).Why this model
17.5% below official
1M context, verified
Three modalities, one model
Much stronger agentic work
Endpoint support
Pricing
Per 1M tokens, pre-discount list price:Specifications
Controlling thinking (the most important section)
Qwen3.8-Max thinks by default, at thexhigh tier. Thinking tokens are billed as output and frequently account for over 90% of it.
Seven values, four real tiers
The parameter accepts 7 values but maps to only 4 real tiers in practice:max does not think harder than xhigh. Any other value returns a 400 listing the legal set.
How to turn thinking off
enable_thinking: false in extra_body and chat_template_kwargs: {"enable_thinking": false} are equivalent and also work.
thinking_budget has no effect
Passing 128 / 512 / 4096 all behave identically to the low tier; the number itself is ignored. Use reasoning_effort instead.
Code examples
Python (OpenAI SDK compatible)
Image input
url to an https://... address.
Video input
{"type": "video", "video": [frame1, frame2, ...]}, which requires 4–8000 frames — fewer than 4 returns a 400.
cURL
Tool calling
Tool calling on the Chat Completions endpoint is fully working: single tool, parallel tools, two-round round-trip, picking 1 out of 20 tools, streaming deltas, andparallel_tool_calls: false all verified.
Structured output
response_format with json_schema held strictly in testing: nested objects, enums, arrays, and additionalProperties: false all took effect, with no extra fields and no Markdown fences.
Context caching
- Hit threshold around 1,024 tokens: an 818-token prefix missed; 1,070 tokens and up hit
- Real multi-turn conversations do hit: appending messages turn by turn hit on every round
- Long documents benefit most: 98.6% cached input at 128K, 99.3% at 32K
Using the Anthropic endpoint
/v1/messages works for code integration, but you must strip thinking blocks before replaying history, otherwise you get a 400 (if content is list. item must be dict and key[type] should in dict).
response_format is silently ignored (force a tool call for structured output), tool_choice only accepts the OpenAI format, images must be base64 (remote URLs return 400), and reasoning_effort has no effect (use thinking: {"type": "disabled"} to turn thinking off).
Parameter compatibility
Best practices
Everyday chat and high-volume calls
reasoning_effort="none" explicitly. Measured latency dropped from ~5 s to 2 s, and output tokens to roughly 1/30.Long documents and codebases
Data extraction
json_schema and disable thinking. Conformance is unaffected.Agents and tool orchestration
/v1/chat/completions. Remember to disable thinking when forcing a tool call.FAQ
Why am I still billed a lot of tokens after setting max_tokens?
Why am I still billed a lot of tokens after setting max_tokens?
max_tokens bounds only the visible answer, not the thinking portion. We measured 1,054 output tokens billed at max_tokens=1. Use reasoning_effort="none" to control cost.Why does tool_choice with a named function return 400?
Why does tool_choice with a named function return 400?
reasoning_effort="none" alongside it.Why can't I reach /v1/responses?
Why can't I reach /v1/responses?
/v1/chat/completions instead.Can I use this model in Claude Code?
Can I use this model in Claude Code?
/v1/messages endpoint rejects history messages containing thinking blocks, and Claude Code replays them verbatim. When calling from your own code, strip those blocks and the endpoint works fine.Why is reasoning_tokens sometimes missing from usage?
Why is reasoning_tokens sometimes missing from usage?
reasoning_tokens or cached_tokens — measured at roughly one third of chat requests. This has been reported upstream for alignment. Keep it in mind if you need exact thinking-cost accounting.Why are video calls so slow?
Why are video calls so slow?
Related
- Qwen3.8-Max playground — send requests directly
- Qwen3.6 series (legacy) — the previous five models
- Qwen3.8-Max launch notes — benchmarks and full write-up
- Model pricing — per-model rates, cache pricing, and available endpoints
- Top-up promotions — stackable discounts