NVIDIA NIM
NVIDIA NIM (Inference Microservices) cung cấp các endpoint API tối ưu hóa cho các mô hình AI mã nguồn mở và doanh nghiệp với 1.000 tín dụng suy luận miễn phí khi đăng ký.
BEVFormer Bird's-Eye-View Perception
NVIDIA BEVFormer bird's-eye-view camera fusion perception model on NIM.
SparseDrive End-to-End Perception
NVIDIA SparseDrive end-to-end perception and motion forecasting on NIM.
StreamPETR 3D Perception
NVIDIA StreamPETR real-time 3D object detection microservice on NIM.
Synthetic Video Detector (SVD)
NVIDIA AI synthetic deepfake video detection microservice on NIM.
Cosmos Transfer 1.0 7B
NVIDIA Cosmos Transfer 1.0 7B high-fidelity physical video transfer on NIM.
Cosmos Transfer 2.5 2B
NVIDIA Cosmos Transfer 2.5 2B physical video style transfer on NIM.
Cosmos 3 Nano Reasoner
NVIDIA Cosmos 3 Nano physical spatial dynamics reasoning model on NIM.
Cosmos 3 Nano World Model
NVIDIA Cosmos 3 Nano physical AI world foundation model on NIM.
PaliGemma 3B Vision-Language
Google's 3B vision-language model for image captioning and OCR on NIM.
DiffusionGemma 26B-A4B IT
Google's Diffusion-Transformer hybrid image synthesis model on NIM.
Nemotron 3 Embed 1B
NVIDIA's 1B dense text embedding model with top MTEB retrieval performance.
Riva Translate 4B Instruct v1.1
NVIDIA Riva 4B high-throughput machine translation v1.1 model on NIM.
Riva Translate 4B Instruct v2
NVIDIA Riva 4B neural machine translation v2 model on NIM.
NVIDIA Active Speaker Detection
NVIDIA audio-visual neural active speaker detection microservice.
NVIDIA Background Noise Removal
NVIDIA real-time background noise removal and echo cancellation.
NVIDIA Studio Voice
NVIDIA studio-quality voice enhancement and audio restoration microservice.
Nemotron VoiceChat Duplex
NVIDIA's full-duplex sub-300ms real-time conversational voice agent on NIM.
Magpie TTS Zero-Shot
NVIDIA's neural zero-shot voice cloning and TTS model on NIM.
Llama Guard 4 12B
Meta's 4th gen multimodal safety & content moderation model on NIM.
Llama 3.1 Nemotron Safety Guard 8B v3
NVIDIA 8B safety guard model for prompt injection & jailbreak detection.
Nemotron 3.5 Content Safety
NVIDIA NeMo Guardrails content safety & moderation model on NIM.
Ising Calibration 1.0 35B-A3B
NVIDIA 35B-A3B MoE reasoning calibration model on NIM.
Ising Calibration 1.5 31B
NVIDIA 31B reasoning calibration and verification model on NIM.
Laguna XS 2.1
Poolside's specialized coding intelligence model on NVIDIA NIM.
Llama 3.2 90B Vision Instruct
Meta's high-capacity 90B multimodal vision model on NVIDIA NIM.
Llama 3.2 11B Vision Instruct
Meta's lightweight 11B multimodal vision model on NVIDIA NIM.
Gemma 4 31B IT
Google's 2026 Gemma 4 31B instruction-tuned model on NVIDIA NIM.
GPT-OSS 20B
OpenAI's fast 20B open-weights model on NVIDIA NIM for code and agent loops.
GPT-OSS 120B
OpenAI's 120B open-weights model hosted on NVIDIA NIM with 1M context.
Mistral Nemotron
Mistral AI and NVIDIA jointly optimized model for structured JSON and reasoning.
Nemotron 3 Nano Omni
NVIDIA 30B-A3B omni multimodal vision-language reasoning model.
Nemotron 3 Nano 30B
NVIDIA ultra-efficient 30B MoE (3B active) low-latency coding model.
Nemotron 3 Super 120B
NVIDIA 120B MoE (12B active) high-throughput reasoning model on NIM.
Nemotron 3 Ultra Free
NVIDIA flagship 550B MoE (55B active) reasoning & code synthesis model via NIM.
MiniMax M3
MiniMax M3 1M MSA Sparse Attention model on NVIDIA NIM — 59% SWE-bench Pro for autonomous coding agents.
Meta Muse Glimmer 30B
Meta's 30B multimodal reasoning model on NVIDIA NIM — text & image input with tool calling and separate thinking.
Nemotron 3.5 Lightning
NVIDIA's 30B-A3B Mamba-2/MoE hybrid architecture with NVFP4 quantization and millisecond latency.
DeepSeek V4 Flash Free
DeepSeek V4 Flash 284B MoE (0731 release) on NVIDIA NIM — ultra-fast coding & tool calling with 1M context.
DeepSeek V4 Pro
DeepSeek V4 Pro 1.65T MoE flagship (0813 release) via NVIDIA NIM — 1M context, 1,000 free credits.
Kimi K3
Moonshot AI 2.8T Kimi K3 MoE API via NVIDIA NIM — 1M context, native vision, free endpoint with 40 RPM limit.
Dùng thử ngay trên trình duyệt
Chọn một mô hình, dán khóa miễn phí của riêng bạn và chạy. Khóa của bạn được gửi một lần để gọi nhà cung cấp và không bao giờ được lưu trên máy chủ của chúng tôi.
Chọn một mô hình, dán khóa miễn phí của riêng bạn và chạy. Khóa của bạn được gửi một lần để gọi nhà cung cấp và không bao giờ được lưu trên máy chủ của chúng tôi.
Phản hồi từ cộng đồng
Bình luận được liên kết với tài khoản GitHub của bạn — đăng nhập để tham gia thảo luận.