Model library

Every model listed here is one NovaServe actually serves — the id shown is the exact id the inference API accepts. Live status and measured latency for each one are on the status page.

Showing 57 of 57 models

anthropic/claude-haiku-4-5

Text

Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.

Fast, low-cost Claude for high-volume tasks.

anthropicchattools
AnthropicChat completions

$8/ 1M output tokens

$1.6 / 1M input tokens

anthropic/claude-opus-4-5

Text

Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.

Anthropic's strongest model for long-horizon agentic and coding work.

anthropicchattools
AnthropicChat completions

$40/ 1M output tokens

$8 / 1M input tokens

anthropic/claude-sonnet-4-5

Text

Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.

Balanced Claude for everyday coding, analysis and tool use.

anthropicchattools
AnthropicChat completions

$24/ 1M output tokens

$4.8 / 1M input tokens

mistral/codestral-latest

Text

Codestral by Mistral, served through the NovaServe gateway.

mistralchattoolsliveness monitored
MistralChat completions

$1.92/ 1M output tokens

$0.64 / 1M input tokens

Use via API

deepseek/deepseek-reasoner

Text

Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.

DeepSeek's chain-of-thought reasoning tier.

deepseekchattools
DeepSeekChat completions

$0.672/ 1M output tokens

$0.448 / 1M input tokens

deepseek/deepseek-chat

Text

Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.

Cost-efficient open-weight general model.

deepseekchattools
DeepSeekChat completions

$0.672/ 1M output tokens

$0.448 / 1M input tokens

mistral/devstral-medium-latest

Text

Devstral Medium by Mistral, served through the NovaServe gateway.

mistralchattoolsliveness monitored
MistralChat completions

$3.2/ 1M output tokens

$0.64 / 1M input tokens

Use via API

google/gemini-2.5-flash

Text

Balanced Gemini with lower cost and latency than Pro.

googlechattoolsliveness monitored
GoogleChat completions

$4/ 1M output tokens

$0.48 / 1M input tokens

Use via API

google/gemini-2.5-flash-lite

Text

Cheapest Gemini tier for simple, high-volume work.

googlechattoolsliveness monitored
GoogleChat completions

$0.64/ 1M output tokens

$0.16 / 1M input tokens

Use via API

google/gemini-2.5-pro

Text

Large-context multimodal reasoning across text, image, audio and video.

googlechattoolsliveness monitored
GoogleChat completions

$16/ 1M output tokens

$2 / 1M input tokens

Use via API

google/gemini-3-flash-preview

Text

Preview Flash model balancing speed and capability.

googlechattoolsliveness monitored
GoogleChat completions

$4/ 1M output tokens

$0.48 / 1M input tokens

Use via API

google/gemini-3.1-flash-lite

Text

Cost-efficient Gemini for classification and summarization.

googlechattoolsliveness monitored
GoogleChat completions

$0.64/ 1M output tokens

$0.16 / 1M input tokens

Use via API

google/gemini-3.1-pro-preview

Text

Google's strongest reasoning model, with very large context.

googlechattoolsliveness monitored
GoogleChat completions

$16/ 1M output tokens

$2 / 1M input tokens

Use via API

google/gemini-3.5-flash

Text

High-efficiency Gemini for fast agentic workflows.

googlechattoolsliveness monitored
GoogleChat completions

$4/ 1M output tokens

$0.48 / 1M input tokens

Use via API

google/gemini-3.6-flash

Text

Latest Gemini Flash — fast coding, reasoning and agent loops.

googlechattoolsliveness monitored
GoogleChat completions

$4/ 1M output tokens

$0.48 / 1M input tokens

Use via API

google/gemini-3.1-flash-tts-preview

Audio

Newest Gemini speech preview.

googletext to speech
GoogleSpeech API

$0.03/ 1K characters

Use via API

google/gemini-2.5-flash-tts

Audio

Gemini speech synthesis with prebuilt voices.

googletext to speech
GoogleSpeech API

$0.03/ 1K characters

Use via API

google/gemini-2.5-flash-lite-preview-tts

Audio

Lightweight Gemini speech for short clips.

googletext to speech
GoogleSpeech API

$0.03/ 1K characters

Use via API

google/gemini-2.5-pro-tts

Audio

Highest-quality Gemini speech synthesis.

googletext to speech
GoogleSpeech API

$0.03/ 1K characters

Use via API

openai/gpt-image-1-mini

Image

Smaller, cheaper OpenAI image model for drafts.

openaiimage generationediting
OpenAIImages API

$0.25/ image

Use via API

openai/gpt-image-2

Image

OpenAI image model — strong at legible text and typography.

openaiimage generationediting
OpenAIImages API

$0.25/ image

Use via API

openai/gpt-5

Text

General-purpose all-rounder with text and image input.

openaichattoolsreasoningliveness monitored
OpenAIResponses API

$16/ 1M output tokens

$2 / 1M input tokens

Use via API

openai/gpt-5-mini

Text

Lower-cost GPT-5 tier for general workloads.

openaichattoolsreasoningliveness monitored
OpenAIResponses API

$3.2/ 1M output tokens

$0.4 / 1M input tokens

Use via API

openai/gpt-5-nano

Text

Fastest GPT-5 tier for simple, high-volume calls.

openaichattoolsreasoningliveness monitored
OpenAIResponses API

$0.64/ 1M output tokens

$0.08 / 1M input tokens

Use via API

openai/gpt-5.2

Text

Prior-generation reasoning model, still strong on multi-step problems.

openaichattoolsreasoningliveness monitored
OpenAIResponses API

$16/ 1M output tokens

$2 / 1M input tokens

Use via API

openai/gpt-5.4

Text

Frontier-class coding and analysis at a lower price point.

openaichattoolsreasoningliveness monitored
OpenAIResponses API

$16/ 1M output tokens

$2 / 1M input tokens

Use via API

openai/gpt-5.4-mini

Text

Strong mini model for sub-agents and high-volume coding work.

openaichattoolsreasoningliveness monitored
OpenAIResponses API

$3.2/ 1M output tokens

$0.4 / 1M input tokens

Use via API

openai/gpt-5.4-nano

Text

Cheapest GPT-5.4 tier — classification, extraction and ranking.

openaichattoolsreasoningliveness monitored
OpenAIResponses API

$0.64/ 1M output tokens

$0.08 / 1M input tokens

Use via API

openai/gpt-5.4-pro

Text

Premium reasoning variant for the most complex tasks.

openaichattoolsreasoning
OpenAIResponses API

$192/ 1M output tokens

$24 / 1M input tokens

Use via API

openai/gpt-5.5

Text

Frontier model for complex coding, analysis and professional work.

openaichattoolsreasoningliveness monitored
OpenAIResponses API

$16/ 1M output tokens

$2 / 1M input tokens

Use via API

openai/gpt-5.5-pro

Text

Extended-reasoning variant for the hardest problems.

openaichattoolsreasoning
OpenAIResponses API

$192/ 1M output tokens

$24 / 1M input tokens

Use via API

openai/gpt-5.6-luna

Text

Fast, low-cost model for short replies and high-volume traffic.

openaichattoolsreasoningliveness monitored
OpenAIResponses API

$1.92/ 1M output tokens

$0.24 / 1M input tokens

Use via API

openai/gpt-5.6-sol

Text

Flagship reasoning model — the default behind RouteNova for hard tasks.

openaichattoolsreasoningliveness monitored
OpenAIResponses API

$16/ 1M output tokens

$2 / 1M input tokens

Use via API

openai/gpt-5.6-terra

Text

Balanced everyday model: near-flagship quality at lower cost.

openaichattoolsreasoningliveness monitored
OpenAIResponses API

$5.76/ 1M output tokens

$0.72 / 1M input tokens

Use via API

groq/openai/gpt-oss-120b

Text

GPT-OSS 120B by Groq, served through the NovaServe gateway.

groqchattoolsliveness monitored
GroqChat completions

$1.2/ 1M output tokens

$0.24 / 1M input tokens

Use via API

groq/openai/gpt-oss-20b

Text

GPT-OSS 20B by Groq, served through the NovaServe gateway.

groqchattoolsliveness monitored
GroqChat completions

$0.8/ 1M output tokens

$0.16 / 1M input tokens

Use via API

xai/grok-4.1

Text

Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.

xAI's frontier reasoning model with real-time knowledge.

xaichattools
xAIChat completions

$24/ 1M output tokens

$4.8 / 1M input tokens

xai/grok-4.1-fast

Text

Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.

Low-latency Grok tier for agent loops and tool calling.

xaichattools
xAIChat completions

$0.8/ 1M output tokens

$0.32 / 1M input tokens

xai/grok-code-fast-1

Text

Awaiting provider funding. This model is wired up but its provider account needs funding before it can be called. Requests are refused until our team tops it up.

Grok tuned for fast code generation and edits.

xaichattools
xAIChat completions

$2.4/ 1M output tokens

$0.32 / 1M input tokens

groq/groq/compound

Text

Groq Compound by Groq, served through the NovaServe gateway.

groqchattoolsliveness monitored
GroqChat completions

$1.2/ 1M output tokens

$0.24 / 1M input tokens

Use via API

groq/groq/compound-mini

Text

Groq Compound Mini by Groq, served through the NovaServe gateway.

groqchattoolsliveness monitored
GroqChat completions

$0.8/ 1M output tokens

$0.16 / 1M input tokens

Use via API

mistral/magistral-small-latest

Text

Magistral Small by Mistral, served through the NovaServe gateway.

mistralchattoolsliveness monitored
MistralChat completions

$2.4/ 1M output tokens

$0.8 / 1M input tokens

Use via API

mistral/ministral-3b-latest

Text

Ministral 3B by Mistral, served through the NovaServe gateway.

mistralchattoolsliveness monitored
MistralChat completions

$0.064/ 1M output tokens

$0.064 / 1M input tokens

Use via API

mistral/ministral-8b-latest

Text

Ministral 8B by Mistral, served through the NovaServe gateway.

mistralchattoolsliveness monitored
MistralChat completions

$0.16/ 1M output tokens

$0.16 / 1M input tokens

Use via API

mistral/mistral-large-latest

Text

Mistral Large by Mistral, served through the NovaServe gateway.

mistralchattoolsliveness monitored
MistralChat completions

$9.6/ 1M output tokens

$3.2 / 1M input tokens

Use via API

mistral/mistral-medium-latest

Text

Mistral Medium by Mistral, served through the NovaServe gateway.

mistralchattoolsliveness monitored
MistralChat completions

$3.2/ 1M output tokens

$0.64 / 1M input tokens

Use via API

mistral/mistral-small-latest

Text

Mistral Small by Mistral, served through the NovaServe gateway.

mistralchattoolsliveness monitored
MistralChat completions

$1.92/ 1M output tokens

$0.64 / 1M input tokens

Use via API

google/gemini-2.5-flash-image

Image

Quick drafts and iterations at the lowest image price.

googleimage generationediting
GoogleImages API

$0.25/ image

Use via API

google/gemini-3.1-flash-image

Image

Fast image generation and editing at pro-level quality.

googleimage generationediting
GoogleImages API

$0.25/ image

Use via API

google/gemini-3-pro-image

Image

Highest-fidelity image generation and editing.

googleimage generationediting
GoogleImages API

$0.25/ image

Use via API

openai/gpt-4o-transcribe

Audio

Accurate speech-to-text with punctuation and formatting.

openaispeech to text
OpenAITranscription API

$0.04/ audio minute

Use via API

openai/gpt-4o-mini-transcribe

Audio

Faster, cheaper transcription for bulk audio.

openaispeech to text
OpenAITranscription API

$0.04/ audio minute

Use via API

openai/gpt-4o-mini-tts

Audio

Text-to-speech with steerable delivery and adjustable speed.

openaitext to speech
OpenAISpeech API

$0.03/ 1K characters

Use via API

groq/qwen/qwen3.6-27b

Text

Qwen 3.6 27B by Groq, served through the NovaServe gateway.

groqchattoolsliveness monitored
GroqChat completions

$0.944/ 1M output tokens

$0.464 / 1M input tokens

Use via API

google/veo-3.1

Video

Highest-quality video generation with audio, up to 8s.

googletext to videowith audio
GoogleVideo jobs API

$0.7/ second of video

Use via API

google/veo-3.1-fast

Video

Faster Veo tier for iteration at moderate cost.

googletext to videowith audio
GoogleVideo jobs API

$0.38/ second of video

Use via API

google/veo-3.1-lite

Video

Cheapest Veo tier — default for drafts and previews.

googletext to videowith audio
GoogleVideo jobs API

$0.22/ second of video

Use via API