Soft orange and teal gradient introducing Orq.ai pricing

Route free. Scale securely.

One AI gateway for every LLM: an OpenAI-compatible API for 500+ models, plus an MCP Gateway for every tool. LLM observability, AI governance and managed agents are built in.

Route free. Scale securely.

One AI gateway for every LLM: an OpenAI-compatible API for 500+ models, plus an MCP Gateway for every tool. LLM observability, AI governance and managed agents are built in.

Start free

Pay-as-you-go

Route every model and every MCP tool through one governed, OpenAI-compatible endpoint, and scale AI features in production.

500+

Models

30+

Providers

1M

Free BYOK req / mo

4.5%

Platform fee

AI Gateway

MCP Gateway

OpenAI & Anthropic compatible API

Smart Router

Budgets and alerts

Unlimited AI Gateway seats

PII detection and redaction

30-day retention

Custom

Enterprise

For organizations with custom requirements and heavy AI workloads.

Everything in Pay-as-you-go, plus

AI Governance

Audit logs

On-prem and VPC deployment options

SOC 2 report, ISO 27001

Custom DPA

Forward Deployed Engineers

Uptime SLA

Slack / Teams support

Custom

Enterprise

For organizations with custom requirements and heavy AI workloads.

Everything in Pay-as-you-go, plus

AI Governance

Audit logs

On-prem and VPC deployment options

SOC 2 report, ISO 27001

Custom DPA

Forward Deployed Engineers

Uptime SLA

Slack / Teams support

Trusted by

Capgemini logo
Vattenfall logo
AFAS Software logo
hear.com
bunq
Moneybird logo
yoco
Capgemini logo
Vattenfall logo
AFAS Software logo
hear.com
bunq
Moneybird logo
yoco
Capgemini logo
Vattenfall logo
AFAS Software logo
hear.com
bunq
Moneybird logo
yoco
Capgemini logo
Vattenfall logo
AFAS Software logo
hear.com
bunq
Moneybird logo
yoco

Every plan includes

The parts you shouldn't have to pay extra for

One gateway, every model & tool

500+ models from 30+ providers behind one OpenAI-compatible API, plus the MCP Gateway for governed tool calls. Fallbacks, retries, caching and load balancing built in.

Observability from request one

Every routed request becomes a trace: cost, tokens, latency and full payloads, with native OpenTelemetry support.

Runtime controls

Budgets and spend limits, guardrail rules and PII redaction enforced on every request; alert or hard-block before the spend happens.

No lock-in

Integrate with all your existing agents and keys without rebuilds, keep the runtime controls and observability on top, and export all your data easily.

What it actually costs

Estimate your Pay-as-you-go bill

Model traffic is billed at provider rates. This estimates the platform side.

To use the AI Gateway, no seats are needed
Orq.ai keysYour own keys
€500
250,000
Enable AI observability
Platform fee 4.5%€22.50
250,000 requests€0
Estimated monthly total€22.50
How is this calculated?
  • Each request through the AI Gateway produces 2 spans and is assumed to average 8,000 tokens across input and output.
  • Processed data is estimated at 4 bytes per token with a 2x storage overhead, so about 64 KB per request.
  • Each agent run is assumed to produce 6 spans: the model calls plus tool, retrieval and evaluation steps. Those spans are added to the spans from your requests, so a run costs both a run and its spans.
  • Estimates exclude optional add-ons, annual commitments and Enterprise volume rates. Real usage varies with agent complexity and payload size.

Full comparison

Compare plans

Full comparison

Compare plans

Full comparison

Compare plans

Swipe to compare plans →

Pay-as-you-go

Enterprise

Compare plans

Pay-as-you-go

Enterprise

Usage & limits

The meters

Orq-managed model fee

Applies when using Orq models and keys

4.5% on credits

Volume terms

Bring Your Own Key

Applies when using your own models and keys

1M requests/month free, 4% after

Included

AI Gateway seats

Your team members who receive API keys and use the AI Gateway

Unlimited

Unlimited

Spans

Every routed request, agent step, evaluator or guardrail run is one span

100k / month

then €7 / 100k

Custom

Processed data

Inputs, outputs, prompts and metadata ingested; resets each billing cycle

1 GB / month

then €3 / GB

Custom

Data retention

30 days

Custom

API rate limits

100 req / min

Custom

AI Studio

Build

AI Studio seats

People building in AI Studio

€35 / seat / month

Custom

Agent runs

One invocation of an agent through the Agent Runtime

500 runs / month

then €0.01 / run

Custom

Knowledge Bases & Agent Memory

Embed, ingest, retrieve, parse and chunk documents

€500 / month

2.5 GB document processing, then €0.10 / MB

Included

Teams

Enterprise SSO (Okta, Microsoft), SAML/OIDC, SSO enforcement, fine-grained RBAC, dedicated Slack channel

€300 / month

Included

AI Gateway

Route

Model catalog

500+ models · 30+ providers

500+ models · 30+ providers

OpenAI and Anthropic compatible API

One base-URL swap, with streaming and structured outputs

Python & Node SDKs

Official orq-ai-sdk and @orq-ai/node clients

MCP Gateway

Connect remote MCP servers; route, govern and observe tool calls with the same policies as model traffic

Skills Gateway

Reusable, versioned instruction packs served to agents and coding assistants

Bring your own keys (BYOK)

Billing and data residency stay with your provider; 1M requests / month free, then 4%

Bring your own model (BYOM)

Connect fine-tuned or self-hosted models via Azure AI Foundry, Vertex AI or any OpenAI-compatible endpoint

Smart Router

Auto-routes each request to the best model by cost/quality strategy; decided in under 40 ms

Fallbacks & retries

Exponential backoff, cross-provider fallback chains

Load Balancing Strategies

Weighted, round-robin and latency-based strategies

Routing rules

Condition-based redirects by header, identity, metadata or project

LLM response caching

Up to ~95% latency reduction on cache hits

Prompt caching

Provider-level prompt caching for reduced token costs

Multimodal endpoints

Image generation & understanding, video, audio, speech, PDFs, embedding, reranking, translation, moderation

Built-in web search

Web search tool built into the Responses API

Coding assistants & frameworks

Claude Code, Cursor, Codex + 25+ agent frameworks

Budgets & spend limits

Per workspace, project, user, API key, provider or model; alert or hard-block

Guardrails & guardrail rules

Block non-compliant requests and responses at the gateway

Plugins

PII redaction and response healing: redact before send, repair malformed JSON

Zero data retention routing

Restrict routing to providers that store nothing

AI Observability

See

Real-time traces & spans

Full span tree per request, visible in ~10 seconds

Cost & token tracking

Per request, model, provider and project

OpenTelemetry ingestion

Native OTLP endpoint; 15+ framework instrumentors

Identities

Attribute cost, tokens and errors to end-users, teams or clients

Dashboards & Reporting API

Latency p50/p95/p99, TTFT, error and pass rates

Threads & sessions

Conversation-, session- and user-level grouping

Online evaluators & live scoring

Alerts & notifiers

Email, Slack and webhook alerts on trace activity

Trace automations

Sample traces into datasets and review queues automatically

SIEM exporter

Stream traces and audit events into Splunk, Datadog or any SIEM

AI Governance

Control

Control Tower

Live oversight of every agent: cost, errors, zombie detection, trust levels

Audit logs

28 entity types, 26 action types

Model governance

Restrict which projects can use which models

Data retention controls & masking

Auto-deletion windows, per-request input/output masking

EU data residency

EU-built, EU-hosted, EU model routing

No training on your data

Prompts, outputs and traces are never used to train models

GDPR

Built-in tools to keep your AI GDPR-compliant: EU-hosted processing, Article 28 DPA, PII redaction, masking and retention controls

EU AI Act

Built-in tools to meet your EU AI Act obligations: traceability, logging, guardrails and human-oversight workflows

Managed Agents

Run

Agent runtime

Single- and multi-agent systems with sub-agents and versioning; trigger via the Responses API, expose via A2A

Tools

Function, HTTP, JSON, Python and remote MCP servers + built-ins

Knowledge bases (RAG)

Parsing, chunking, hybrid search, reranking, agentic RAG, RAG evaluators & chunk explorer

Document processing priority

Files are parsed and indexed ahead of the standard queue

Memory stores

Long-term memory that persists across sessions

Prompt management & playgrounds

Versioned prompts, templating, side-by-side model comparison

Environments

Route develop, staging and production versions with environment tags

Experiments, evaluators & datasets

Batch evals with cost/latency metrics, via UI or CI/CD

Evaluator library

40+ prebuilt evaluators in the Hub: 16 function, 10 LLM, 12 RAGAS

Annotations & review queues

Structured human feedback, written back to traces

Agent schedules

Cron-based recurring runs

Agent templates

Share and clone across the workspace

AI Chat

One workspace chat for every model and agent

Exports

Traces, experiments and datasets as CSV / JSON

Agent sandbox (on-prem)

Isolated pods for customer Python in self-hosted installs

Administration

Enterprise SSO (SAML / OIDC)

Okta, Microsoft Entra, any IdP, with enforcement

Teams add-on

Projects & team access

Project containers with team-level access and per-project API keys

API & management keys

Scoped data-plane and admin-plane keys with rotation

Webhooks

Subscribe to agent, deployment, prompt and LLM events over HTTP

CLI

Manage every resource from the terminal and CI; 40+ command groups

Terraform

Manage gateway, prompt and agent configuration declaratively as infra-as-code

MCP

Operate Orq from Claude Code, Cursor and 24 more coding assistants

Service & Support

Uptime SLA

Support center

Email support with priority SLA

Slack / Teams channel

Teams add-on

Dedicated solutions engineer

SOC 2 Type II & ISO 27001 certified

The platform is certified on every plan

SOC 2 report, HIPAA BAA & custom DPA

Custom contracts & redlines

Compare plans

Pay-as-you-go

Enterprise

Usage & limits

The meters

Orq-managed model fee

Applies when using Orq models and keys

4.5% on credits

Volume terms

Bring Your Own Key

Applies when using your own models and keys

1M requests/month free, 4% after

Included

AI Gateway seats

Your team members who receive API keys and use the AI Gateway

Unlimited

Unlimited

Spans

Every routed request, agent step, evaluator or guardrail run is one span

100k / month

then €7 / 100k

Custom

Processed data

Inputs, outputs, prompts and metadata ingested; resets each billing cycle

1 GB / month

then €3 / GB

Custom

Data retention

30 days

Custom

API rate limits

100 req / min

Custom

AI Studio

Build

AI Studio seats

People building in AI Studio

€35 / seat / month

Custom

Agent runs

One invocation of an agent through the Agent Runtime

500 runs / month

then €0.01 / run

Custom

Knowledge Bases & Agent Memory

Embed, ingest, retrieve, parse and chunk documents

€500 / month

2.5 GB document processing, then €0.10 / MB

Included

Teams

Enterprise SSO (Okta, Microsoft), SAML/OIDC, SSO enforcement, fine-grained RBAC, dedicated Slack channel

€300 / month

Included

AI Gateway

Route

Model catalog

500+ models · 30+ providers

500+ models · 30+ providers

OpenAI and Anthropic compatible API

One base-URL swap, with streaming and structured outputs

Python & Node SDKs

Official orq-ai-sdk and @orq-ai/node clients

MCP Gateway

Connect remote MCP servers; route, govern and observe tool calls with the same policies as model traffic

Skills Gateway

Reusable, versioned instruction packs served to agents and coding assistants

Bring your own keys (BYOK)

Billing and data residency stay with your provider; 1M requests / month free, then 4%

Bring your own model (BYOM)

Connect fine-tuned or self-hosted models via Azure AI Foundry, Vertex AI or any OpenAI-compatible endpoint

Smart Router

Auto-routes each request to the best model by cost/quality strategy; decided in under 40 ms

Fallbacks & retries

Exponential backoff, cross-provider fallback chains

Load Balancing Strategies

Weighted, round-robin and latency-based strategies

Routing rules

Condition-based redirects by header, identity, metadata or project

LLM response caching

Up to ~95% latency reduction on cache hits

Prompt caching

Provider-level prompt caching for reduced token costs

Multimodal endpoints

Image generation & understanding, video, audio, speech, PDFs, embedding, reranking, translation, moderation

Built-in web search

Web search tool built into the Responses API

Coding assistants & frameworks

Claude Code, Cursor, Codex + 25+ agent frameworks

Budgets & spend limits

Per workspace, project, user, API key, provider or model; alert or hard-block

Guardrails & guardrail rules

Block non-compliant requests and responses at the gateway

Plugins

PII redaction and response healing: redact before send, repair malformed JSON

Zero data retention routing

Restrict routing to providers that store nothing

AI Observability

See

Real-time traces & spans

Full span tree per request, visible in ~10 seconds

Cost & token tracking

Per request, model, provider and project

OpenTelemetry ingestion

Native OTLP endpoint; 15+ framework instrumentors

Identities

Attribute cost, tokens and errors to end-users, teams or clients

Dashboards & Reporting API

Latency p50/p95/p99, TTFT, error and pass rates

Threads & sessions

Conversation-, session- and user-level grouping

Online evaluators & live scoring

Alerts & notifiers

Email, Slack and webhook alerts on trace activity

Trace automations

Sample traces into datasets and review queues automatically

SIEM exporter

Stream traces and audit events into Splunk, Datadog or any SIEM

AI Governance

Control

Control Tower

Live oversight of every agent: cost, errors, zombie detection, trust levels

Audit logs

28 entity types, 26 action types

Model governance

Restrict which projects can use which models

Data retention controls & masking

Auto-deletion windows, per-request input/output masking

EU data residency

EU-built, EU-hosted, EU model routing

No training on your data

Prompts, outputs and traces are never used to train models

GDPR

Built-in tools to keep your AI GDPR-compliant: EU-hosted processing, Article 28 DPA, PII redaction, masking and retention controls

EU AI Act

Built-in tools to meet your EU AI Act obligations: traceability, logging, guardrails and human-oversight workflows

Managed Agents

Run

Agent runtime

Single- and multi-agent systems with sub-agents and versioning; trigger via the Responses API, expose via A2A

Tools

Function, HTTP, JSON, Python and remote MCP servers + built-ins

Knowledge bases (RAG)

Parsing, chunking, hybrid search, reranking, agentic RAG, RAG evaluators & chunk explorer

Document processing priority

Files are parsed and indexed ahead of the standard queue

Memory stores

Long-term memory that persists across sessions

Prompt management & playgrounds

Versioned prompts, templating, side-by-side model comparison

Environments

Route develop, staging and production versions with environment tags

Experiments, evaluators & datasets

Batch evals with cost/latency metrics, via UI or CI/CD

Evaluator library

40+ prebuilt evaluators in the Hub: 16 function, 10 LLM, 12 RAGAS

Annotations & review queues

Structured human feedback, written back to traces

Agent schedules

Cron-based recurring runs

Agent templates

Share and clone across the workspace

AI Chat

One workspace chat for every model and agent

Exports

Traces, experiments and datasets as CSV / JSON

Agent sandbox (on-prem)

Isolated pods for customer Python in self-hosted installs

Administration

Enterprise SSO (SAML / OIDC)

Okta, Microsoft Entra, any IdP, with enforcement

Teams add-on

Projects & team access

Project containers with team-level access and per-project API keys

API & management keys

Scoped data-plane and admin-plane keys with rotation

Webhooks

Subscribe to agent, deployment, prompt and LLM events over HTTP

CLI

Manage every resource from the terminal and CI; 40+ command groups

Terraform

Manage gateway, prompt and agent configuration declaratively as infra-as-code

MCP

Operate Orq from Claude Code, Cursor and 24 more coding assistants

Service & Support

Uptime SLA

Support center

Email support with priority SLA

Slack / Teams channel

Teams add-on

Dedicated solutions engineer

SOC 2 Type II & ISO 27001 certified

The platform is certified on every plan

SOC 2 report, HIPAA BAA & custom DPA

Custom contracts & redlines

FAQ

Frequently asked questions

Usage & Billing
What is a span, and what counts as one?

A span is the unit of observability and billing in Orq.ai. Each discrete operation in your AI application creates one: a model call routed through the AI Gateway, a step inside an agent run (a tool call, reasoning step or retrieval), an evaluator or guardrail execution, or an external span you send via the SDK or OpenTelemetry. A trace is the complete end-to-end interaction; spans are its steps. An agent run is one invocation of an agent and typically contains multiple spans. Billing is based on spans.

What’s the difference between AI Gateway seats and AI Studio seats?

An AI Gateway seat is a team member who holds an API key and routes traffic through the Gateway: engineers shipping features, data scientists running jobs, or anyone calling models from their own code and tools. Gateway seats are unlimited on every plan and never billed per person. What you pay for is their usage, metered as spans and processed data, and you can attribute and budget it per person via Identities. An AI Studio seat is a person working in the Orq.ai interface: writing prompts, configuring agents, running experiments and reading dashboards. Studio seats are the only per-person charge, at €35 per seat per month on Pay-as-you-go. So a developer who only needs an API key costs you nothing per seat. Add a Studio seat when they need the visual workspace.

What happens when I exceed my included usage?

On Pay-as-you-go, nothing breaks: your traffic keeps routing, and usage beyond the included allowance is billed at the metered rates shown in the table as separate line items on your invoice. You stay in control before that happens: set budgets with alerts, or hard caps per workspace, project, user or API key. Until you add billing details, requests are rate-limited (50 per day) rather than billed.

When am I billed, and how does seat proration work?

Billing runs at the end of each monthly cycle, and included allowances reset at the start of the next. Seats adjust automatically: inviting a member mid-cycle charges a prorated amount immediately; removing one frees the seat at the start of the next cycle.

How can I reduce my bill?

Four levers, all included:

Gateway & Models
Do you mark up model traffic?

No. Models are billed at provider list prices; Orq.ai adds no markup. With BYOK, billing stays entirely with your provider: routing is free up to 1M requests a month, then a 4% fee applies. If you route on Orq credits instead, a 4.5% platform fee applies on top-ups.

Can I use the AI Gateway without the rest of the platform?

Yes. The Gateway works standalone: swap your base URL, keep your OpenAI-compatible code, and you’re routing in minutes. Observability is captured from the first request, and evals, agents and governance are already there when you need them. No migration, ever.

What latency does the Gateway add?

Smart Router classifies each request in under 40 ms before it reaches a model, and cached responses return up to ~95% faster than a live call. Fallbacks and retries only add latency when a provider actually fails.

Can I use my own private or fine-tuned models?

Yes, on every plan. Connect fine-tuned, self-hosted or private models through Azure AI Foundry, Google Vertex AI or LiteLLM; they join the same catalog as the 500+ public models and work with routing, fallbacks and observability like any other model.

Plans & Subscription
Which plan is right for me?

Pay-as-you-go: almost everyone: start free with the included allowances (100k spans, 1 GB data, 500 agent runs), then pay only for what you use. Every feature is unlocked from day one. Enterprise: organizations that need custom volumes, enterprise security (SSO, RBAC, audit logs), on-prem or VPC deployment, and dedicated support.

Can I change or cancel my subscription?

Yes, upgrade, downgrade or cancel Pay-as-you-go any time, self-serve. Changes take effect at the end of the current billing cycle. Enterprise agreements run on annual terms.

Do you offer discounts for startups or academia?

Yes, we offer startup-friendly terms and academic arrangements case by case. Talk to us.

Security & Deployment
Where does my data live, and is Orq.ai compliant?

Orq.ai is EU-built and EU-hosted, SOC 2 Type II certified, GDPR-compliant with a standard DPA, and aligned with the EU AI Act. Zero-data-retention routing can restrict traffic to providers that store nothing. Full details, reports and controls are available in the Orq.ai Trust Center at Orq.ai Trust Center (trust.orq.ai).

Do you train models on my data?

No. Your prompts, outputs and traces are never used to train models, ours or anyone else’s. You control retention with configurable auto-deletion, and can mask inputs and outputs per request.

Can we self-host or deploy in our VPC?

Yes, Enterprise customers can deploy in their own VPC on AWS or Azure via the marketplaces, or fully on-premise with a single Helm chart, including air-gapped environments. Custom contracts and redlines are available on annual Enterprise agreements.

Future-proof solution

Why teams switch

One control tower across teams

Unite engineering, product, and data teams in one place. Shared truth, role-based workflows, and human-in-the-loop feedback that drives continuous improvement.

Deploy anywhere, safely

Our cloud, your cloud, or your servers. Private connections supported. Roll out safely and roll back fast.

Compliant, secure and flexible

SOC 2-certified, GDPR-compliant, and aligned with the EU AI Act. Manage risk responsibly with EU or US data residency across open and closed ecosystems.

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.