Skip to main content
APIYI supports 400+ mainstream AI models. This page provides detailed model information, pricing, and usage instructions.
Enterprise-grade Professional and Stable AI Large Model API Hub All models are officially sourced and forwarded, with ~20% off pricing (combining top-up bonuses and exchange rate advantages), aggregating various excellent large models. No speed limits, no expiration, no account ban risks, pay-as-you-go billing, long-term reliable service.
The following are currently stably supplied popular models. For the complete model list and real-time pricing, visit the APIYI Console Pricing Page or the on-site Model Pricing Overview.
Model Upgrade Recommendations: We recommend using the latest models for best performance, but please note:
  1. Initial instability is common: Newly launched models may experience slow responses, timeouts, or occasional errors due to limited compute capacity at the vendor — this typically stabilizes within days to weeks
  2. Check parameter compatibility: New models may introduce or change parameters (e.g., max_completion_tokens replacing max_tokens). Before upgrading, verify that your API parameters remain compatible with older models
  3. Always test before going live: Before deploying a new model to production, thoroughly validate it in a test environment to ensure output quality and API compatibility meet expectations
A ”—” in the context column means the vendor has not published a definitive figure; refer to the console and official documentation. All prices are list prices per 1M tokens and can be combined with top-up bonuses.

Model Categories

🤖 OpenAI Series

🆕 Latest Models

GPT-5.6 dual-group supply: the default official-relay groups (default / svip) match official pricing, with up to 20% top-up bonus (~83% of list). The CodexReverse group carries a default 0.7 discount, better for cost-sensitive batch workloads. Existing GPT-5.5 code only needs the model field swapped. See the GPT-5.6 launch notes.
GPT Pro series (e.g. gpt-5.5-pro, gpt-5.4-pro) usage notes:
  1. /v1/responses endpoint only: cannot use /v1/chat/completions — switch your SDK / code to the Responses API before calling
  2. Very expensive: a single call may cost several dollars, and it is open to the SVIP group only to prevent accidental use on the Default group
  3. Not recommended for non-professional needs: use GPT-5.6 Sol / Terra for everyday tasks; Pro suits only research and top-tier work demanding extreme reasoning depth

✅ Stable / Classic

GPT-5 and newer series usage notes:
  1. Temperature parameter temperature must be set to 1 (only supports 1)
  2. Use max_completion_tokens instead of max_tokens
  3. Do not pass top_p parameter
Image and Video Generation Models have been moved to a dedicated page. Visit Image & Video Generation Models for the full list and pricing.

🎭 Claude Series (Anthropic)

🆕 Latest Models

30-day data retention compliance for Fable 5 / Mythos 5: because of these models’ advanced capabilities, Bedrock and Anthropic run additional checks for high-risk misuse. For these two models, inputs and outputs are retained for 30 days to detect serious abuse — accessed by automated safety systems by default, with human review only when a potential harm is flagged. Retention happens on the AWS / Anthropic side; APIYI is a pure transparent proxy and retains nothing itself. Other Claude models are unaffected. See the Fable 5 launch notes.

✅ Stable / Classic

Latest: Claude Opus 5 approaches Fable 5 intelligence at half the price ($5/$25, same as Opus 4.8) — a painless upgrade. Sonnet 5 is billed as the most agentic Sonnet yet, close to Opus 4.8 but faster and cheaper. Stable: Opus 4.7 and Sonnet 4.6 are battle-tested for production workflows already in place; Haiku 4.5 offers 2x speed at great value.

🌟 Google Gemini Series

🆕 Latest Models

Note: Gemini 3 Pro Preview was discontinued on March 9, 2026. Please migrate to Gemini 3.1 Pro Preview or Gemini 3.6 Flash.

✅ Stable / Classic

Latest: Gemini 3.6 Flash and 3.5 Flash-Lite were validated end to end with 30+ test cases each on both the native Gemini format and the OpenAI-compatible format, with all native tools enabled. Read the 3.6 Flash overview and 3.5 Flash-Lite overview before integrating. Stable: Gemini 2.5 Pro (2M context) and Gemini 2.5 Flash suit settled production environments.

🚀 xAI Grok Series

🆕 Latest Models

✅ Stable / Classic

Grok 4.5 token efficiency: on SWE Bench Pro it averages just 15,954 output tokens, about 1/4.2 of Opus 4.8 (max), halving task steps and cutting real spend on the same job. Note that requests above 200K tokens are billed at the vendor’s higher long-context rate.

🔍 DeepSeek Series

🐘 Chinese Model Series

Zhipu AI (GLM)

🆕 Latest: GLM-5.2 | ✅ Stable / Classic: GLM-5.1, GLM-5, GLM-4.6
GLM-5.2 Features:
  • Context jumps from GLM-5.1’s 200K to 1M tokens — load project-scale code and multiple long documents at once
  • 744B MoE, Terminal-Bench 2.1 81.0 / SWE-bench Pro 62.1, leading open-source on long-horizon coding
  • Direct official Alibaba Cloud relay with pricing aligned to the vendor (converted at the fixed 1:7 rate); ~85% of list with top-up bonuses
  • For long tasks, raise your timeout above 600 seconds. See the GLM-5.2 launch notes

Alibaba Qwen

🆕 Latest: Qwen3.7-Max | ✅ Stable / Classic: Qwen3.6 series, Qwen Max / Plus / Turbo

Moonshot Kimi Series

🆕 Latest: Kimi K3 | ✅ Stable / Classic: Kimi K2.6, K2.5, K2
Kimi K3 usage notes:
  • Exactly 1,048,576 tokens of context with flat pricing across the whole range — drop in whole repos or long documents without worrying about tier jumps
  • Max output defaults to 131,072 and can be raised to 1,048,576 tokens; thinking runs at full tilt, so budget output tokens accordingly
  • Automatic context caching bills cached input at $0.30/1M (1/10 of the uncached rate), a big win for multi-turn conversations

ByteDance Seed Series

🆕 Latest: Seed 2.1 Turbo | ✅ Stable / Classic: Seed 2.0 Pro / Lite / Mini
Seed 2.1 Turbo enables deep thinking by default: with no parameters set, even a one-line question emits hundreds of thinking tokens first (a one-sentence self-introduction measured 444 output tokens, 409 of them thinking), and thinking is billed as normal output tokens. For high-frequency short Q&A, make thinking: {"type": "disabled"} your baseline. See the integration guide.

🌐 MiniMax Series

🆕 Latest: MiniMax-M3 | ✅ Stable / Classic: MiniMax-M2.7, M2.5
MiniMax-M3 billing notes:
  • Tiered billing: $0.30 input / $1.20 output per 1M tokens for the 0-512K range; anything beyond is billed at a higher tier
  • Model names are case-sensitive: MiniMax-M3 (as are siblings such as MiniMax-M2.7-highspeed)
  • Estimate costs before running million-token context jobs

💰 Pricing Information

Billing Methods

  • Pay-as-you-go: Charged based on actual Token usage
  • No minimum charge: Use what you pay for, balance never expires
  • Real-time deduction: Fees deducted from balance immediately after each call

Pricing Advantages

  • Official source forwarding with slight price advantages
  • Bulk users can contact customer service for better pricing
  • New users get 3 million tokens testing credit upon registration

View Real-time Pricing

Visit the APIYI Console Pricing Page for the latest pricing on all models, or the on-site Model Pricing Overview.

🛠️ Usage Recommendations

Model Selection Guide

Programming Development
  • Top performance: Claude Opus 5 (near Fable 5 intelligence at half the price), GPT-5.6 Sol (Terminal-Bench 2.1 88.8%), Claude Sonnet 5 (SWE-bench 85.2%), Kimi K3
  • High cost-performance: GPT-5.6 Terra (GPT-5.5-level performance at half price), Gemini 3.6 Flash, GLM-5.2, MiniMax-M3, DeepSeek V4 Flash
  • Alternatives: Grok 4.5, Qwen3.7-Max, Kimi K2.6, Gemini 3.5 Flash
Text Creation
  • Top choice: GPT-5.6 Sol / Terra, Claude Opus 5, Claude Sonnet 5, Gemini 3.6 Flash
  • Alternatives: chat-latest, Claude Sonnet 4.6, GLM-5.2, GPT-4.1
Quick Response
  • Top choice: Gemini 3.5 Flash-Lite (zero thinking by default, ~2s), Claude Haiku 4.5 (2x faster), GPT-5.6 Luna
  • Alternatives: Gemini 3.1 Flash Lite, Gemini 2.5 Flash, Seed 2.1 Turbo (disable thinking explicitly)
Long Text Processing
  • Ultra-long context: Gemini 2.5 Pro (2M), Kimi K3 (1M, flat pricing), MiniMax-M3 (1M), GLM-5.2 (1M), Claude Opus 5 (1M)
  • Note: some models switch to a higher billing tier past a threshold (Grok 4.5 above 200K, MiniMax-M3 above 512K) — estimate costs before long jobs
Image Generation
  • Latest recommendation: gpt-image-2 (native 4K, precise size/quality control), Nano Banana Pro (4K HD, best text rendering)
  • High cost-performance: Nano Banana 2 ($0.055/image, from $0.025 pay-as-you-go), Nano Banana Lite (~4s per image, $0.025/image)
  • Professional design: Seedream 5.0 Pro ($0.12/call, interactive editing + up to 10 reference images), Seedream 5.0 Lite ($0.035/image)
  • Cheapest: gpt-image-2-all (reverse channel, $0.03/image)
Video Generation
  • Official relay first choice: VEO 3.1 Official ($0.3 / $1.2 per call, 4/6/8s, native synced audio)
  • Chinese-vendor workhorses: Seedance 2.0 series (standard / fast / mini), Wan2.7, HappyHorse 1.1
  • Note: the Sora 2 channels have been retired — use the options above. See Image & Video Generation Models
Web Search
  • Native web access: Grok 4 All, Grok 3 All (no tool call needed)
  • Search grounding: Gemini 3.6 Flash / 3.5 Flash-Lite (native tools enabled)

Cost Optimization Recommendations

  1. Tiered Usage: Use cheaper models for simple tasks, advanced models for complex tasks
  2. Test Optimization: Test with small models first, use large models after determining needs
  3. Batch Processing: Choose Luna / Lite / Mini tiers for large volumes of similar tasks
  4. Cache Reuse: Lean on each vendor’s context cache (Kimi K3, Seed 2.1 Turbo and the Claude series bill cached input as low as 1/10)
  5. Turn off unneeded thinking: some models think by default (Seed 2.1 Turbo, Gemini 3.6 Flash) — disabling it for short high-frequency Q&A saves both money and time
Model list is continuously updated. We will promptly add newly released excellent models. For the latest launches, follow the Changelog. For specific model needs or bulk requirements, please contact customer service.