Model catalog
One endpoint, every model. Compare context windows, capabilities and real per-token pricing side by side, then swap the model string in your request — nothing else changes.
- 351
- Models
- 35
- Providers
- 11
- Free to try
- $0.0001
- Cheapest input / 1M
7 of 351 models
- Language
Nemotron 3.5 Lightning 30B
nvidia/nemotron-3.5-lightning
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA. The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total.
- Context
- 1M
- In / 1M
- Free
- Out / 1M
- Free
- Free
- Reasoning
- Tool use
- +1
- Language
Nemotron 3.5 Lightning 30B (Free)
nvidia/nemotron-3.5-lightning-free
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA. The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total.
- Context
- 1M
- In / 1M
- Free
- Out / 1M
- Free
- Free
- Reasoning
- Tool use
- +1
- Language
Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b
A 550B parameter (55B active) open reasoning model from NVIDIA, built for long-running agent workflows. It uses a hybrid Mamba-Transformer MoE architecture and supports a 1M token context window.
- Context
- 1M
- In / 1M
- $0.600
- Out / 1M
- $2.40
- Reasoning
- Tool use
- Implicit caching
- Language
Nvidia Nemotron Nano 12B V2 VL
nvidia/nemotron-nano-12b-v2-vl
The model is an auto-regressive vision language model that uses an optimized transformer architecture. The model enables multi-image reasoning and video understanding, along with strong document intelligence, visual Q&A and summarization capabilities.
- Context
- 131K
- In / 1M
- $0.200
- Out / 1M
- $0.600
- Reasoning
- Tool use
- Vision
- Language
NVIDIA Nemotron 3 Super 120B A12B
nvidia/nemotron-3-super-120b-a12b
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. It delivers up to 7x higher throughput, providing fast, cost-efficient inference for agentic tasks. Additionally, a long context window gives the model long-term memory, preventing AI agents from losing focus on long, multi-step tasks and ensuring high-accuracy results. Fully open with weights, datasets, and recipes, Super allows easy customization and secure deployment anywhere.
- Context
- 256K
- In / 1M
- $0.150
- Out / 1M
- $0.650
- Reasoning
- Tool use
- Language
Nemotron 3 Nano 30B A3B
nvidia/nemotron-3-nano-30b-a3b
NVIDIA Nemotron 3 Nano is an open reasoning model optimized for fast, cost-efficient inference. Built with a hybrid MoE and Mamba architecture and trained on NVIDIA-curated synthetic reasoning data, it delivers strong multi-step reasoning with stable latency and predictable performance for agentic and production workloads.
- Context
- 262K
- In / 1M
- $0.050
- Out / 1M
- $0.240
- Reasoning
- Tool use
- Language
Nvidia Nemotron Nano 9B V2
nvidia/nemotron-nano-9b-v2
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so.\
- Context
- 131K
- In / 1M
- $0.060
- Out / 1M
- $0.230
- Reasoning
- Tool use
