Gemini 3.5 Flash Lite
GoogleReleased Jul 21, 2026
- providers
- 1
- context
- 1.0M
- max out
- 65.5K
- input /M
- $0.30
- output /M
- $2.50
- cached input /M
- $0.03
- cache write /M
- $0.30
- reasoning /M
- $2.50
About
Drafted from the provider's own page.
Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing. The model supports text, image, video, audio, and PDF inputs, and is designed for high-volume agentic workflows, simple data extraction, and applications where latency and API cost are the primary constraints.
- knowledge cutoff
- Mar 1, 2026
- open weights
- no
- best for
- extractioncodingcompute
Capabilities
What the model accepts and returns, where a probe through our own gateway proved it.
- streaming
- tools
- json mode
- vision
- audio
Providers
What you pay through us for each provider, and its own list price where that differs.
Google
google-aistudio/gemini-3.5-flash-lite- input /M
- $0.30
- output /M
- $2.50
- cache read /M
- $0.03
- cache write /M
- $0.30
Cache read $0.03/M · Cache write $0.30/M
cache read list price /M$0.03- Web search
- $0.014 per request
The provider's list price. We do not charge for searches yet.
1.0M context65.5K max out
Speed
- fastest over 24 hours
- 1.0k t/s
Fastest: Google
- Google616 ms to first token1.0k t/smedian of 24
Served
- served in the last 24 hours
- 100.00%
Across 1 provider, 25 of 25 requests
- Google100.00% · 25 of 25
Rebuilt once a day from the requests we sent, so it is a daily figure and not a live one.
Not yet. These figures come from Artificial Analysis, which the publish pipeline reads once a day.
Integrations
What the model holds, rather than a compatibility grade: nothing here measures how well a particular tool gets on with it, so nothing here scores one.
- answers on
- chat completions
- accepts
- textimagevideoaudiopdf
- returns
- text
- reasoning
- response_format
- tool_choice
- tools
- web_search
Use this model
Works with the OpenAI SDK. Point it at our base URL and use this id.
curl https://api.hopscotchlabs.ai/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google-aistudio/gemini-3.5-flash-lite",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 256
}'