Every model listed here is one NovaServe actually serves — the id shown is the exact id the inference API accepts. Live status and measured latency for each one are on the status page.
Showing 57 of 57 models
anthropic/claude-haiku-4-5
Text
Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.
Fast, low-cost Claude for high-volume tasks.
anthropicchattools
AnthropicChat completions
$8/ 1M output tokens
$1.6 / 1M input tokens
anthropic/claude-opus-4-5
Text
Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.
Anthropic's strongest model for long-horizon agentic and coding work.
anthropicchattools
AnthropicChat completions
$40/ 1M output tokens
$8 / 1M input tokens
anthropic/claude-sonnet-4-5
Text
Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.
Balanced Claude for everyday coding, analysis and tool use.
anthropicchattools
AnthropicChat completions
$24/ 1M output tokens
$4.8 / 1M input tokens
mistral/codestral-latest
Text
Codestral by Mistral, served through the NovaServe gateway.
Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.
DeepSeek's chain-of-thought reasoning tier.
deepseekchattools
DeepSeekChat completions
$0.672/ 1M output tokens
$0.448 / 1M input tokens
deepseek/deepseek-chat
Text
Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.
Cost-efficient open-weight general model.
deepseekchattools
DeepSeekChat completions
$0.672/ 1M output tokens
$0.448 / 1M input tokens
mistral/devstral-medium-latest
Text
Devstral Medium by Mistral, served through the NovaServe gateway.
Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.
xAI's frontier reasoning model with real-time knowledge.
xaichattools
xAIChat completions
$24/ 1M output tokens
$4.8 / 1M input tokens
xai/grok-4.1-fast
Text
Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.
Low-latency Grok tier for agent loops and tool calling.
xaichattools
xAIChat completions
$0.8/ 1M output tokens
$0.32 / 1M input tokens
xai/grok-code-fast-1
Text
Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.
Grok tuned for fast code generation and edits.
xaichattools
xAIChat completions
$2.4/ 1M output tokens
$0.32 / 1M input tokens
groq/groq/compound
Text
Groq Compound by Groq, served through the NovaServe gateway.