
Start free
Pay-as-you-go
Route every model and every MCP tool through one governed, OpenAI-compatible endpoint, and scale AI features in production.
500+
Models
30+
Providers
1M
Free BYOK req / mo
4.5%
Platform fee
AI Gateway
MCP Gateway
OpenAI & Anthropic compatible API
Smart Router
Budgets and alerts
Unlimited AI Gateway seats
PII detection and redaction
30-day retention
Trusted by
Every plan includes
The parts you shouldn't have to pay extra for
One gateway, every model & tool
500+ models from 30+ providers behind one OpenAI-compatible API, plus the MCP Gateway for governed tool calls. Fallbacks, retries, caching and load balancing built in.
Observability from request one
Every routed request becomes a trace: cost, tokens, latency and full payloads, with native OpenTelemetry support.
Runtime controls
Budgets and spend limits, guardrail rules and PII redaction enforced on every request; alert or hard-block before the spend happens.
No lock-in
Integrate with all your existing agents and keys without rebuilds, keep the runtime controls and observability on top, and export all your data easily.
What it actually costs
Estimate your Pay-as-you-go bill
Model traffic is billed at provider rates. This estimates the platform side.
| Platform fee 4.5% | €22.50 |
| 250,000 requests | €0 |
How is this calculated?
- Each request through the AI Gateway produces 2 spans and is assumed to average 8,000 tokens across input and output.
- Processed data is estimated at 4 bytes per token with a 2x storage overhead, so about 64 KB per request.
- Each agent run is assumed to produce 6 spans: the model calls plus tool, retrieval and evaluation steps. Those spans are added to the spans from your requests, so a run costs both a run and its spans.
- Estimates exclude optional add-ons, annual commitments and Enterprise volume rates. Real usage varies with agent complexity and payload size.
FAQ
Frequently asked questions
What is a span, and what counts as one?
A span is the unit of observability and billing in Orq.ai. Each discrete operation in your AI application creates one: a model call routed through the AI Gateway, a step inside an agent run (a tool call, reasoning step or retrieval), an evaluator or guardrail execution, or an external span you send via the SDK or OpenTelemetry. A trace is the complete end-to-end interaction; spans are its steps. An agent run is one invocation of an agent and typically contains multiple spans. Billing is based on spans.
What’s the difference between AI Gateway seats and AI Studio seats?
An AI Gateway seat is a team member who holds an API key and routes traffic through the Gateway: engineers shipping features, data scientists running jobs, or anyone calling models from their own code and tools. Gateway seats are unlimited on every plan and never billed per person. What you pay for is their usage, metered as spans and processed data, and you can attribute and budget it per person via Identities. An AI Studio seat is a person working in the Orq.ai interface: writing prompts, configuring agents, running experiments and reading dashboards. Studio seats are the only per-person charge, at €35 per seat per month on Pay-as-you-go. So a developer who only needs an API key costs you nothing per seat. Add a Studio seat when they need the visual workspace.
What happens when I exceed my included usage?
On Pay-as-you-go, nothing breaks: your traffic keeps routing, and usage beyond the included allowance is billed at the metered rates shown in the table as separate line items on your invoice. You stay in control before that happens: set budgets with alerts, or hard caps per workspace, project, user or API key. Until you add billing details, requests are rate-limited (50 per day) rather than billed.
When am I billed, and how does seat proration work?
Billing runs at the end of each monthly cycle, and included allowances reset at the start of the next. Seats adjust automatically: inviting a member mid-cycle charges a prorated amount immediately; removing one frees the seat at the start of the next cycle.
How can I reduce my bill?
Four levers, all included:
Do you mark up model traffic?
No. Models are billed at provider list prices; Orq.ai adds no markup. With BYOK, billing stays entirely with your provider: routing is free up to 1M requests a month, then a 4% fee applies. If you route on Orq credits instead, a 4.5% platform fee applies on top-ups.
Can I use the AI Gateway without the rest of the platform?
Yes. The Gateway works standalone: swap your base URL, keep your OpenAI-compatible code, and you’re routing in minutes. Observability is captured from the first request, and evals, agents and governance are already there when you need them. No migration, ever.
What latency does the Gateway add?
Smart Router classifies each request in under 40 ms before it reaches a model, and cached responses return up to ~95% faster than a live call. Fallbacks and retries only add latency when a provider actually fails.
Can I use my own private or fine-tuned models?
Yes, on every plan. Connect fine-tuned, self-hosted or private models through Azure AI Foundry, Google Vertex AI or LiteLLM; they join the same catalog as the 500+ public models and work with routing, fallbacks and observability like any other model.
Which plan is right for me?
Pay-as-you-go: almost everyone: start free with the included allowances (100k spans, 1 GB data, 500 agent runs), then pay only for what you use. Every feature is unlocked from day one. Enterprise: organizations that need custom volumes, enterprise security (SSO, RBAC, audit logs), on-prem or VPC deployment, and dedicated support.
Can I change or cancel my subscription?
Yes, upgrade, downgrade or cancel Pay-as-you-go any time, self-serve. Changes take effect at the end of the current billing cycle. Enterprise agreements run on annual terms.
Do you offer discounts for startups or academia?
Yes, we offer startup-friendly terms and academic arrangements case by case. Talk to us.
Where does my data live, and is Orq.ai compliant?
Orq.ai is EU-built and EU-hosted, SOC 2 Type II certified, GDPR-compliant with a standard DPA, and aligned with the EU AI Act. Zero-data-retention routing can restrict traffic to providers that store nothing. Full details, reports and controls are available in the Orq.ai Trust Center at Orq.ai Trust Center (trust.orq.ai).
Do you train models on my data?
No. Your prompts, outputs and traces are never used to train models, ours or anyone else’s. You control retention with configurable auto-deletion, and can mask inputs and outputs per request.
Can we self-host or deploy in our VPC?
Yes, Enterprise customers can deploy in their own VPC on AWS or Azure via the marketplaces, or fully on-premise with a single Helm chart, including air-gapped environments. Custom contracts and redlines are available on annual Enterprise agreements.
Future-proof solution
Why teams switch
One control tower across teams
Unite engineering, product, and data teams in one place. Shared truth, role-based workflows, and human-in-the-loop feedback that drives continuous improvement.
Deploy anywhere, safely
Our cloud, your cloud, or your servers. Private connections supported. Roll out safely and roll back fast.
Compliant, secure and flexible
SOC 2-certified, GDPR-compliant, and aligned with the EU AI Act. Manage risk responsibly with EU or US data residency across open and closed ecosystems.
