Model
Model explorer

Gemini 3.5 Flash-Lite

CLOSED
Google · Gemini 3.5 family · released Jul 21, 2026

Google DeepMind's fastest and least-expensive model in the Gemini 3.5 family, released 2026-07-21 alongside Gemini 3.6 Flash. A low-latency, high-throughput reasoning model (thinking supported despite the Lite tier) built as a 'subagent' for high-volume automation -- agentic search, document processing, classification and routing. 1M-token context, natively multimodal (text/image/audio/video in, text out), ~350 output tokens/sec (official). Model ID gemini-3.5-flash-lite, generally available on the Gemini API/AI Studio/Vertex; $0.30/M input, $2.50/M output (cache-hit input ~$0.03/M). Benchmarks are Google self-reported from the joint launch blog, which shows large jumps over the prior Gemini 3.1 Flash-Lite (e.g. GDPval-AA v2 1140 vs 642, Terminal-Bench 2.1 54% vs 31%). No independent per-benchmark scores exist yet (Artificial Analysis publishes only a composite Intelligence Index of 36).

ReasoningCodingVisionFunction callingTool useAgentic
2116.0
Elo · rank #51
Parameters
Undisclosed
Active params
Undisclosed
Context
1M tokens
Architecture
Sparse Mixture-of-Experts Transformer (configuration undisclosed); reasoning-capable multimodal Lite model in the Gemini 3.5 series
License
Proprietary
Languages
API price (in/out)
$0.3 / $2.5
Modalities
text · vision · audio · video
Benchmark results
Bar shows position within the tracked field; marker = field best
GDPval-AAAgents1140#31
best: Claude Fable 5 · 1932
OSWorld-VerifiedAgents74.0%#11
best: Claude Fable 5 · 85.0%
SWE-bench ProCoding54.2%#32
best: Claude Fable 5 · 80.0%
Terminal-Bench 2.0Coding54.0%#45
best: GPT-5.6 · 88.8%
Run it locally
Closed weights — available via API only. No local deployment.
Input / M tok
$0.3
Output / M tok
$2.5
Gemini 3.5 family
Elo progression across releases
API price $0.3/$2.5 · each benchmark row carries its own source badge (see methodology)