Text Generation APIs
LLMs for text completion and chat
Gemini 3.7 Flash
Official Google AI Studio Gemini 3.7 Flash API with hybrid thinking, 2M context, and generous daily free quotas.
GLM 5.3 Flash
Vultr Serverless Inference endpoint for GLM 5.3 Flash ($0.10 in, $0.35 out).
MiniCPM 5 1B
Ultra-fast 1B parameter on-device language model served with high throughput on AMD GPU Cloud.
DeepSeek V4 Flash Vision (Exp)
Official AMD endpoint for DeepSeek V4 Flash Vision (Exp) multimodal reasoning with 1M context.
GPT-OSS 20b
OpenAI's fast 20B open-weights model on NVIDIA NIM for code and agent loops.
GPT-OSS 120b
OpenAI's 120B open-weights model hosted on NVIDIA NIM with 1M context.
QwQ 32b
Free QwQ 32b API via Cloudflare Workers AI with 10,000 daily free Neurons. Zero setup fee, no credit card required.
Qwen3 30b A3b FP8
Free Qwen3 30b A3b FP8 API via Cloudflare Workers AI with 10,000 daily free Neurons. Zero setup fee, no credit card required.
Mistral Small 3.1 24b Instruct
Free Mistral Small 3.1 24b Instruct API via Cloudflare Workers AI with 10,000 daily free Neurons. Zero setup fee, no credit card required.
Mistral 7b Instruct V0.2 LoRA
Free Mistral 7b Instruct V0.2 LoRA API via Cloudflare Workers AI with 10,000 daily free Neurons. Zero setup fee, no credit card required.
Llama Guard 3 8b
Free Llama Guard 3 8b API via Cloudflare Workers AI with 10,000 daily free Neurons. Zero setup fee, no credit card required.
Llama 4 Scout 17b 16e Instruct
Free Llama 4 Scout 17b 16e Instruct API via Cloudflare Workers AI with 10,000 daily free Neurons. Zero setup fee, no credit card required.