Gemini 3.1 Flash-Lite
Google's efficiency tier and a long-term stable target, priced for workloads measured in millions of calls rather than thousands. It gives up the deeper reasoning of 3.5 Flash, so it suits mechanical tasks — classification, extraction, moderation — where a million tokens of context still helps but chain-of-thought does not.
Pricing checked against provider documentation on . How we verify
gemini-3.1-flash-lite Input
$0.25
/ 1M tokens
Output
$1.50
/ 1M tokens
Context
1.05M
1,048,576 tokens
Max output
66K
tokens per request
What it costs in practice
A workload of 10M input and 2M output tokens per month — roughly a moderately busy production assistant — runs $5.50/month on Gemini 3.1 Flash-Lite. Routing that same volume through the batch API brings it to $2.75. The next cheaper option, GPT-5.6 Luna , would cost $4.40.
Where Gemini 3.1 Flash-Lite fits
- Content moderation at volume
- Classification and routing
- Lightweight summarization
- Cheap long-context retrieval
Specifications
- API model ID
- gemini-3.1-flash-lite
- Provider
- Model family
- Gemini 3.1
- Released
- March 2026
- Knowledge cutoff
- January 2025
- Context window
- 1,048,576 tokens
- Max output
- 65,536 tokens
- Native reasoning
- No
- Relative latency
- low
- Input modalities
- text, image, audio, video, pdf
- Output modalities
- text
- Open weights
- No
- Cached input
- —
- Batch discount
- 50%
Frequently asked
How much does Gemini 3.1 Flash-Lite cost?
Gemini 3.1 Flash-Lite costs $0.25 per million input tokens and $1.50 per million output tokens. Batch processing is 50% cheaper.
What is Gemini 3.1 Flash-Lite's context window?
Gemini 3.1 Flash-Lite accepts up to 1,048,576 tokens of context and can generate up to 65,536 output tokens per request.
Is Gemini 3.1 Flash-Lite a reasoning model?
No. Gemini 3.1 Flash-Lite answers directly without a native thinking phase, which keeps latency and cost predictable.
What is Gemini 3.1 Flash-Lite best for?
Ultra-cheap high-volume tasks. Google's efficiency tier and a long-term stable target, priced for workloads measured in millions of calls rather than thousands. It gives up the deeper reasoning of 3.5 Flash, so it suits mechanical tasks — classification, extraction, moderation — where a million tokens of context still helps but chain-of-thought does not.