Skip to main content
The cost-efficiency workhorse of the line: multimodal reasoning with 1M context and 131K max output, suited to visual coding, document work and long-video analysis at volume.

Specifications

Pricing

Prices in USD per 1M tokens ($/1M).
The table shows list prices. Recharge promotions and group discounts stack; the console reflects the actual charge in real time. See Pricing and Recharge promotions.

Endpoints

Billing groups

Some groups carry extra discounts, stackable with recharge bonuses. See Tokens and groups.

Supported features

Example request

The example below calls qwen3.8-flash through OpenAI Chat Completions (/v1/chat/completions). Point base_url at https://api.apiyi.com/v1 — everything else matches the official API.
Read the key from an environment variable, never hard-coded. In production, issue separate tokens per use case so you can revoke and attribute usage individually.

Qwen3.7-Flash

Details for a model in the same family

Qwen3.8-27B

Details for a model in the same family

Model pricing directory

Live pricing, endpoints and groups for all 289 models
Specs on this page are maintained by hand in models/data/model-details.json; pricing and endpoints come from the live pricing API, updated 2026-09-03 01:13 (UTC+8), specs last verified 2026-09-03.