Skip to main content
ai-query supports multiple AI providers through a unified interface. Each provider has its own factory function that creates a LanguageModel instance.

Prompt caching

Prompt caching is automatic. There is no ai-query cache configuration and no application-level response cache: every request still runs the model and produces a new response.
  • OpenAI, DeepSeek, and supported Gemini models use their provider defaults.
  • Anthropic requests opt into its automatic ephemeral prompt cache.
  • Other providers retain their existing behavior.
Keep repeated content at the beginning of the prompt to improve the chance of a provider cache hit. Stable system instructions, tool definitions, and earlier conversation turns should come before request-specific content. Avoid timestamps, request IDs, or frequently changing tool schemas in the shared prefix. Inspect result.usage.cached_tokens, cache_write_tokens, and cache_miss_tokens to see the cache information reported by the provider. A zero means absent or unreported; ai-query does not infer cache activity that the provider does not expose.

Supported Providers

OpenAI

GPT-5, GPT-4o, and o-series reasoning models

Anthropic

Claude 3.5 Sonnet, Claude 3 Opus, and more

Google

Gemini 2.0, Gemini 1.5 Pro and Flash

OpenRouter

Access any model through one API

DeepSeek

DeepSeek Chat and Reasoner models

Groq

Groq LPU Inference

Workers AI

Cloudflare Workers AI via OpenAI-compatible endpoints

Custom Providers

Build your own provider adapters and wrappers

Meta Llama

Official Meta Llama API

xAI

Grok Models via xAI API

Cloud Providers

Bedrock

AWS Bedrock - managed foundation models
Bedrock requires an extra dependency: pip install ai-query[bedrock] or uv add ai-query[bedrock]

Quick Start

Each provider requires an API key set as an environment variable:

Usage

All providers follow the same pattern - import the factory function and use it with generate_text() or stream_text():

Switching Providers

One of the key benefits of ai-query is how easy it is to switch between providers:

Model input capabilities

Declare input modalities when conversation history can contain images or files:
Before each provider request, ai-query projects the canonical conversation history into the target model’s capabilities without mutating stored messages:
  • Native endpoints such as Gemini 3 keep rich tool results inline.
  • Vision models on chat-compatible endpoints receive tool-result images in a follow-up multimodal user message.
  • Text-only models receive explicit omission placeholders instead of image bytes.
Leave input_modalities unspecified to retain provider behavior when capabilities are not known. Provider catalogs and applications should set this metadata whenever it is available.

Provider Options

Each provider supports specific options through the provider_options parameter:
See individual provider pages for available options.

Deterministic Tests with the Faux Provider

The public faux provider runs the normal generation, streaming, tool, and agent loops without network requests or credentials. Queue one response per model step:
Use a response factory when the next scripted response should inspect the model, messages, tools, or provider options captured in FauxCall. Queue exhaustion is an error by design, and assert_exhausted() catches unused scripted steps.

OpenAI-Compatible Providers

OpenRouter, DeepSeek, Groq, and Workers AI use OpenAI-compatible APIs internally. This means:
  • They share the same request/response format
  • Provider options still use their own provider keys ("openrouter", "deepseek", "groq", "workers_ai")
  • Tool calling works identically