Skip to content

Gemini 3.1 Flash-Lite

Google's efficiency tier and a long-term stable target, priced for workloads measured in millions of calls rather than thousands. It gives up the deeper reasoning of 3.5 Flash, so it suits mechanical tasks — classification, extraction, moderation — where a million tokens of context still helps but chain-of-thought does not.

Pricing checked against provider documentation on . How we verify

stable Google Gemini 3.1 gemini-3.1-flash-lite

Input

$0.25

/ 1M tokens

Output

$1.50

/ 1M tokens

Context

1.05M

1,048,576 tokens

Max output

66K

tokens per request

What it costs in practice

A workload of 10M input and 2M output tokens per month — roughly a moderately busy production assistant — runs $5.50/month on Gemini 3.1 Flash-Lite. Routing that same volume through the batch API brings it to $2.75. The next cheaper option, GPT-5.6 Luna , would cost $4.40.

Run your own numbers in the calculator →

Where Gemini 3.1 Flash-Lite fits

  • Content moderation at volume
  • Classification and routing
  • Lightweight summarization
  • Cheap long-context retrieval

Specifications

API model ID
gemini-3.1-flash-lite
Provider
Google
Model family
Gemini 3.1
Released
March 2026
Knowledge cutoff
January 2025
Context window
1,048,576 tokens
Max output
65,536 tokens
Native reasoning
No
Relative latency
low
Input modalities
text, image, audio, video, pdf
Output modalities
text
Open weights
No
Cached input
Batch discount
50%

Frequently asked

How much does Gemini 3.1 Flash-Lite cost?

Gemini 3.1 Flash-Lite costs $0.25 per million input tokens and $1.50 per million output tokens. Batch processing is 50% cheaper.

What is Gemini 3.1 Flash-Lite's context window?

Gemini 3.1 Flash-Lite accepts up to 1,048,576 tokens of context and can generate up to 65,536 output tokens per request.

Is Gemini 3.1 Flash-Lite a reasoning model?

No. Gemini 3.1 Flash-Lite answers directly without a native thinking phase, which keeps latency and cost predictable.

What is Gemini 3.1 Flash-Lite best for?

Ultra-cheap high-volume tasks. Google's efficiency tier and a long-term stable target, priced for workloads measured in millions of calls rather than thousands. It gives up the deeper reasoning of 3.5 Flash, so it suits mechanical tasks — classification, extraction, moderation — where a million tokens of context still helps but chain-of-thought does not.