Code Generation APIs
Models optimized for programming
Nemotron 3 Nano 30B
NVIDIA ultra-efficient 30B MoE (3B active) low-latency coding model.
Nemotron 3 Super 120B
NVIDIA 120B MoE (12B active) high-throughput reasoning model on NIM.
Nemotron 3 Ultra 550B
NVIDIA flagship 550B MoE (55B active) reasoning & code synthesis model via NIM.
Nemotron 3.5 Lightning
NVIDIA's 30B-A3B Mamba-2/MoE hybrid architecture with NVFP4 quantization and millisecond latency.
DeepSeek V4 Flash
DeepSeek V4 Flash ultra-fast MoE coding model free tier on OpenCode Zen with sub-second latency.
Kimi K3
Moonshot AI 2.8T Kimi K3 MoE API via NVIDIA NIM — 1M context, native vision, free endpoint with 40 RPM limit.
MiniMax M3
MiniMax M3 1M MSA Sparse Attention model on NVIDIA NIM — 59% SWE-bench Pro for autonomous coding agents.
MiniMax M3
GMI Cloud 官方提供的 MiniMax M3 推理端点。限时 14 天 100% 全量免费,支持 1M MSA 稀疏注意力与 OpenAI 协议直连。
DeepSeek V4 Pro
Access deepseek-v4-pro free through Vercel's $5 monthly credit.