Skip to main content
The first open-weight model combining frontier agentic coding, a million-token context and native multimodality; MSA sparse attention cuts long-context inference cost to about 1/20 of the previous generation.

Specifications

Pricing

Prices in USD per 1M tokens ($/1M).
The table shows list prices. Recharge promotions and group discounts stack; the console reflects the actual charge in real time. See Pricing and Recharge promotions.

Tiered pricing

This model is tier-priced by the token size of each request (output price = tier input price × output multiplier):
  • 0 – 524,288 tokens: $0.3/1M input
  • Above 524,288 tokens: $0.6/1M input
(already discounted to 50% of the vendor’s list price)

Endpoints

Billing groups

Some groups carry extra discounts, stackable with recharge bonuses. See Tokens and groups.

Supported features

Example request

The example below calls MiniMax-M3 through OpenAI Chat Completions (/v1/chat/completions). Point base_url at https://api.apiyi.com/v1 — everything else matches the official API.
Read the key from an environment variable, never hard-coded. In production, issue separate tokens per use case so you can revoke and attribute usage individually.

Launch announcement

Background, benchmarks and migration notes for MiniMax-M3

Model pricing directory

Live pricing, endpoints and groups for all 308 models
Specs on this page are maintained by hand in models/data/model-details.json; pricing and endpoints come from the live pricing API, updated 2026-07-31 01:30 (UTC+8), specs last verified 2026-07-31.