All models1 shown
Nemotron 3 Ultra 550B A55B
NVIDIAfireworks-ai/nemotron-3-ultra-nvfp4
Released Jun 2, 2026
- providers
- 1
- context
- 262.1K
- max out
- 32.7K
- input /M
- $0.60
- output /M
- $2.40
- cached input /M
- $0.119
- cache write /M
- $0.60
- reasoning /M
- $2.40
also answers tonvidia-nemotron-3-ultra-550b-a55bnvidia/Nemotron-3-Ultra-550b-a55bnvidia/nemotron-3-ultra-550b-a55b-20260604
About
- open weights
- yes
- best for
- extractioncodingcompute
Capabilities
What the model accepts and returns, where a probe through our own gateway proved it.
- streaming
- tools
- json mode
- vision
- audio
Providers
What you pay through us for each provider, and its own list price where that differs.
Fireworks AI
fireworks-ai/nemotron-3-ultra-nvfp4- input /M
- $0.60
- output /M
- $2.40
- cache read /M
- $0.119
- cache write /M
- $0.60
Cache read $0.119/M · Cache write $0.60/M
cache read list price /M$0.12262.1K context32.7K max out
Speed
- fastest over 24 hours
- 178.0 t/s
Fastest: Fireworks AI
- Fireworks AI1,085 ms to first token178.0 t/smedian of 23
Served
- served in the last 24 hours
- 96.15%
Across 1 provider, 25 of 26 requests
- Fireworks AI96.15% · 25 of 26
Rebuilt once a day from the requests we sent, so it is a daily figure and not a live one.
not yet
Not yet. These figures come from Artificial Analysis, which the publish pipeline reads once a day.
Integrations
What the model holds, rather than a compatibility grade: nothing here measures how well a particular tool gets on with it, so nothing here scores one.
- answers on
- chat completions
- accepts
- text
- returns
- text
request parameters it takes
- reasoning
- response_format
- tool_choice
- tools
Use this model
Works with the OpenAI SDK. Point it at our base URL and use this id.
curl https://api.hopscotchlabs.ai/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "fireworks-ai/nemotron-3-ultra-nvfp4",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 256
}'