NVIDIA NIM
O NVIDIA NIM (Microserviços de Inferência) fornece endpoints de API otimizados para modelos corporativos e de código aberto com 1.000 créditos de inferência gratuitos no cadastro.
BEVFormer Bird's-Eye-View Perception
NVIDIA BEVFormer bird's-eye-view camera fusion perception model on NIM.
SparseDrive End-to-End Perception
NVIDIA SparseDrive end-to-end perception and motion forecasting on NIM.
StreamPETR 3D Perception
NVIDIA StreamPETR real-time 3D object detection microservice on NIM.
Synthetic Video Detector (SVD)
NVIDIA AI synthetic deepfake video detection microservice on NIM.
Cosmos Transfer 1.0 7B
NVIDIA Cosmos Transfer 1.0 7B high-fidelity physical video transfer on NIM.
Cosmos Transfer 2.5 2B
NVIDIA Cosmos Transfer 2.5 2B physical video style transfer on NIM.
Cosmos 3 Nano Reasoner
NVIDIA Cosmos 3 Nano physical spatial dynamics reasoning model on NIM.
Cosmos 3 Nano World Model
NVIDIA Cosmos 3 Nano physical AI world foundation model on NIM.
PaliGemma 3B Vision-Language
Google's 3B vision-language model for image captioning and OCR on NIM.
DiffusionGemma 26B-A4B IT
Google's Diffusion-Transformer hybrid image synthesis model on NIM.
Nemotron 3 Embed 1B
NVIDIA's 1B dense text embedding model with top MTEB retrieval performance.
Riva Translate 4B Instruct v1.1
NVIDIA Riva 4B high-throughput machine translation v1.1 model on NIM.
Riva Translate 4B Instruct v2
NVIDIA Riva 4B neural machine translation v2 model on NIM.
NVIDIA Active Speaker Detection
NVIDIA audio-visual neural active speaker detection microservice.
NVIDIA Background Noise Removal
NVIDIA real-time background noise removal and echo cancellation.
NVIDIA Studio Voice
NVIDIA studio-quality voice enhancement and audio restoration microservice.
Nemotron VoiceChat Duplex
NVIDIA's full-duplex sub-300ms real-time conversational voice agent on NIM.
Magpie TTS Zero-Shot
NVIDIA's neural zero-shot voice cloning and TTS model on NIM.
Llama Guard 4 12B
Meta's 4th gen multimodal safety & content moderation model on NIM.
Llama 3.1 Nemotron Safety Guard 8B v3
NVIDIA 8B safety guard model for prompt injection & jailbreak detection.
Nemotron 3.5 Content Safety
NVIDIA NeMo Guardrails content safety & moderation model on NIM.
Ising Calibration 1.0 35B-A3B
NVIDIA 35B-A3B MoE reasoning calibration model on NIM.
Ising Calibration 1.5 31B
NVIDIA 31B reasoning calibration and verification model on NIM.
Laguna XS 2.1
Poolside's specialized coding intelligence model on NVIDIA NIM.
Llama 3.2 90B Vision Instruct
Meta's high-capacity 90B multimodal vision model on NVIDIA NIM.
Llama 3.2 11B Vision Instruct
Meta's lightweight 11B multimodal vision model on NVIDIA NIM.
Gemma 4 31B IT
Google's 2026 Gemma 4 31B instruction-tuned model on NVIDIA NIM.
GPT-OSS 20B
OpenAI's fast 20B open-weights model on NVIDIA NIM for code and agent loops.
GPT-OSS 120B
OpenAI's 120B open-weights model hosted on NVIDIA NIM with 1M context.
Mistral Nemotron
Mistral AI and NVIDIA jointly optimized model for structured JSON and reasoning.
Nemotron 3 Nano Omni
NVIDIA 30B-A3B omni multimodal vision-language reasoning model.
Nemotron 3 Nano 30B
NVIDIA ultra-efficient 30B MoE (3B active) low-latency coding model.
Nemotron 3 Super 120B
NVIDIA 120B MoE (12B active) high-throughput reasoning model on NIM.
Nemotron 3 Ultra Free
NVIDIA flagship 550B MoE (55B active) reasoning & code synthesis model via NIM.
MiniMax M3
MiniMax M3 1M MSA Sparse Attention model on NVIDIA NIM — 59% SWE-bench Pro for autonomous coding agents.
Meta Muse Glimmer 30B
Meta's 30B multimodal reasoning model on NVIDIA NIM — text & image input with tool calling and separate thinking.
Nemotron 3.5 Lightning
NVIDIA's 30B-A3B Mamba-2/MoE hybrid architecture with NVFP4 quantization and millisecond latency.
DeepSeek V4 Flash Free
DeepSeek V4 Flash 284B MoE (0731 release) on NVIDIA NIM — ultra-fast coding & tool calling with 1M context.
DeepSeek V4 Pro
DeepSeek V4 Pro 1.65T MoE flagship (0813 release) via NVIDIA NIM — 1M context, 1,000 free credits.
Kimi K3
Moonshot AI 2.8T Kimi K3 MoE API via NVIDIA NIM — 1M context, native vision, free endpoint with 40 RPM limit.
Experimente no navegador
Escolha um modelo, cole sua própria chave gratuita e execute. Sua chave é enviada uma vez para chamar o provedor e nunca é armazenada em nossos servidores.
Escolha um modelo, cole sua própria chave gratuita e execute. Sua chave é enviada uma vez para chamar o provedor e nunca é armazenada em nossos servidores.
Feedback da comunidade
Os comentários estão vinculados à sua conta GitHub — entre para participar da discussão.