Skip to main content
A 284B-total / 13B-active MoE with 1M context, built for high concurrency and low latency; implicit caching needs no setup and lands near-full hits from the second turn.

Specifications

Pricing

Prices in USD per 1M tokens ($/1M).
The table shows list prices. Recharge promotions and group discounts stack; the console reflects the actual charge in real time. See Pricing and Recharge promotions.

Endpoints

Billing groups

Some groups carry extra discounts, stackable with recharge bonuses. See Tokens and groups.

Supported features

Example request

The example below calls deepseek-v4-flash through OpenAI Chat Completions (/v1/chat/completions). Point base_url at https://api.apiyi.com/v1 — everything else matches the official API.
Read the key from an environment variable, never hard-coded. In production, issue separate tokens per use case so you can revoke and attribute usage individually.

Launch announcement

Background, benchmarks and migration notes for DeepSeek V4 Flash

Overview

Parameters, usage and best practices

DeepSeek V4 Pro

Details for a model in the same family

Model pricing directory

Live pricing, endpoints and groups for all 296 models
Specs on this page are maintained by hand in models/data/model-details.json; pricing and endpoints come from the live pricing API, updated 2026-08-17 13:47 (UTC+8), specs last verified 2026-08-17.