Text Generation APIs
LLMs for text completion and chat
Nemotron 3 Embed 1B
NVIDIA's 1B dense text embedding model with top MTEB retrieval performance.
Riva Translate 4B Instruct v1.1
NVIDIA Riva 4B high-throughput machine translation v1.1 model on NIM.
Riva Translate 4B Instruct v2
NVIDIA Riva 4B neural machine translation v2 model on NIM.
Nemotron VoiceChat Duplex
NVIDIA's full-duplex sub-300ms real-time conversational voice agent on NIM.
Llama Guard 4 12B
Meta's 4th gen multimodal safety & content moderation model on NIM.
Llama 3.1 Nemotron Safety Guard 8B v3
NVIDIA 8B safety guard model for prompt injection & jailbreak detection.
Nemotron 3.5 Content Safety
NVIDIA NeMo Guardrails content safety & moderation model on NIM.
Ising Calibration 1.0 35B-A3B
NVIDIA 35B-A3B MoE reasoning calibration model on NIM.
Ising Calibration 1.5 31B
NVIDIA 31B reasoning calibration and verification model on NIM.
Llama 3.2 90B Vision Instruct
Meta's high-capacity 90B multimodal vision model on NVIDIA NIM.
Llama 3.2 11B Vision Instruct
Meta's lightweight 11B multimodal vision model on NVIDIA NIM.
Gemma 4 31B IT
Google's 2026 Gemma 4 31B instruction-tuned model on NVIDIA NIM.