LanguageModel instance.
Prompt caching
Prompt caching is automatic. There is no ai-query cache configuration and no application-level response cache: every request still runs the model and produces a new response.- OpenAI, DeepSeek, and supported Gemini models use their provider defaults.
- Anthropic requests opt into its automatic ephemeral prompt cache.
- Other providers retain their existing behavior.
result.usage.cached_tokens, cache_write_tokens, and
cache_miss_tokens to see the cache information reported by the provider. A zero
means absent or unreported; ai-query does not infer cache activity that the provider
does not expose.
Supported Providers
OpenAI
GPT-5, GPT-4o, and o-series reasoning models
Anthropic
Claude 3.5 Sonnet, Claude 3 Opus, and more
Gemini 2.0, Gemini 1.5 Pro and Flash
OpenRouter
Access any model through one API
DeepSeek
DeepSeek Chat and Reasoner models
Groq
Groq LPU Inference
Workers AI
Cloudflare Workers AI via OpenAI-compatible endpoints
Custom Providers
Build your own provider adapters and wrappers
Meta Llama
Official Meta Llama API
xAI
Grok Models via xAI API
Cloud Providers
Bedrock
AWS Bedrock - managed foundation models
Quick Start
Each provider requires an API key set as an environment variable:Usage
All providers follow the same pattern - import the factory function and use it withgenerate_text() or stream_text():
Switching Providers
One of the key benefits of ai-query is how easy it is to switch between providers:Model input capabilities
Declare input modalities when conversation history can contain images or files:- Native endpoints such as Gemini 3 keep rich tool results inline.
- Vision models on chat-compatible endpoints receive tool-result images in a follow-up multimodal user message.
- Text-only models receive explicit omission placeholders instead of image bytes.
input_modalities unspecified to retain provider behavior when capabilities
are not known. Provider catalogs and applications should set this metadata whenever
it is available.
Provider Options
Each provider supports specific options through theprovider_options parameter:
Deterministic Tests with the Faux Provider
The publicfaux provider runs the normal generation, streaming, tool, and agent
loops without network requests or credentials. Queue one response per model step:
FauxCall. Queue exhaustion is an
error by design, and assert_exhausted() catches unused scripted steps.
OpenAI-Compatible Providers
OpenRouter, DeepSeek, Groq, and Workers AI use OpenAI-compatible APIs internally. This means:- They share the same request/response format
- Provider options still use their own provider keys (
"openrouter","deepseek","groq","workers_ai") - Tool calling works identically