All models

Gemini 3.5 Flash Lite

Google
google-aistudio/gemini-3.5-flash-lite

Released Jul 21, 2026

providers
1
context
1.0M
max out
65.5K
input /M
$0.30
output /M
$2.50
cached input /M
$0.03
cache write /M
$0.30
reasoning /M
$2.50
also answers togemini-3-5-flash-litegoogle/gemini-3.5-flash-lite-20260721

About

Drafted from the provider's own page.

Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing. The model supports text, image, video, audio, and PDF inputs, and is designed for high-volume agentic workflows, simple data extraction, and applications where latency and API cost are the primary constraints.

knowledge cutoff
Mar 1, 2026
open weights
no
best for
extractioncodingcompute

Capabilities

What the model accepts and returns, where a probe through our own gateway proved it.

  • streaming
  • tools
  • json mode
  • vision
  • audio