Skip to main content
A low-cost, high-throughput tier with adjustable thinking levels, suited to concurrency and batch processing.

Specifications

Pricing

Prices in USD per 1M tokens ($/1M).
The table shows list prices. Recharge promotions and group discounts stack; the console reflects the actual charge in real time. See Pricing and Recharge promotions.

Endpoints

Billing groups

Some groups carry extra discounts, stackable with recharge bonuses. See Tokens and groups.

Supported features

Variants

Example request

The example below calls gemini-3.1-flash-lite through OpenAI Chat Completions (/v1/chat/completions). Point base_url at https://api.apiyi.com/v1 — everything else matches the official API.
Read the key from an environment variable, never hard-coded. In production, issue separate tokens per use case so you can revoke and attribute usage individually.

Launch announcement

Background, benchmarks and migration notes for Gemini 3.1 Flash-Lite

Native Calls

Parameters, usage and best practices

Gemini 3.5 Flash-Lite

Details for a model in the same family

Model pricing directory

Live pricing, endpoints and groups for all 308 models
Specs on this page are maintained by hand in models/data/model-details.json; pricing and endpoints come from the live pricing API, updated 2026-07-31 01:30 (UTC+8), specs last verified 2026-07-31.