AI monitoring · cost tracking · model routing
See every AI call. Control what it costs. Pick the right model.
CEEZ Compass sits between your apps and AI providers like OpenAI and Anthropic. Sign in to watch live traffic, see every dollar by team and model, and decide which model answers which request.
Works with any OpenAI-compatible app · 38 ready-made providers · No account? Request access
What do you want to do?
Start with the question, not the feature.
Is AI working right now?
Live health, speed and failures across every provider.
See how →Where is the money going?
Spend by team, app and model, with a forecast and savings advice.
See how →Can a cheaper model do this?
Rules that send each request to the right model, tested before they go live.
See how →How do I stop surprise bills?
Budgets and limits per team and key, and alerts before you hit them.
See how →What is leaving the building?
Guardrails that mask or block sensitive data on the way out.
See how →How does a team get access?
Self-service keys with limits, owned by a person, shown once.
See how →Monitoring
Know the moment AI misbehaves.
Every request through Compass is recorded: how long it took, whether it worked, which provider answered and what it cost. You get a live picture instead of hearing about it from a customer.
- Live call log: each request with its status, tokens, cost and the reason it went where it did
- Health and latency: success rate, typical and slow answers, time to first word, by model
- Reliability: failovers, retries and rate limits, and which provider is struggling
- Alerts: email, Slack or webhook when errors or budgets cross a line
The benefit: a provider outage becomes a line in a log, because Compass has already moved traffic to your backup.
last hour
p50
p95
today
Recent events
Illustrative.
of $6,000 budget
by month end
14% of spend
all costed
Spend by team
Suggestion
Data team's nightly summaries use your strongest model for short, simple text. A smaller model handled similar requests at about one fifth of the cost.
Turn into a routing rule →
Illustrative figures.
Tokenomics
Know what AI costs, and who spent it.
AI is billed in tokens. Compass counts them on every call and turns them into dollars by team, app, key and model, using the prices you set, including the cheaper rate for cached input.
- Spend and usage over time, with a forecast to month end
- Models and callers: which model and which app cost the most
- Budgets and limits by day or month, per key, project and team: refuse at the limit, or switch to a cheaper model
- Savings advice from your own traffic: repeated questions, over-qualified models
- Chargeback reports per team, ready for finance
The benefit: one provider bill becomes a statement per team, and the next saving is a suggestion, not a hunt.
Model routing
Send each request to the right model.
You rarely need the most expensive model for every question. Write simple rules once, and Compass applies them to every request, with no change to your apps.
- By what the request is: short or long, simple or complex, which tools it uses
- By what is happening: a provider failing, slow, or a team's budget used up
- By policy: keep certain data in certain regions, or on certain providers
- Cheap first, escalate if needed, or split traffic to compare two models
- Test before you commit: replay last week's traffic, or run a rule in shadow mode
The benefit: pay for the strong model only when a request needs it, and never be stuck on one provider.
- IF the request is short and simple → use the fast, low-cost model
- IF the prompt is very long → use the large-context model
- IF it mentions personal data → use only the EU-hosted provider
- IF the chosen provider is failing → use the fallback
- IF the team's budget is used up → downgrade to the cheaper model
- ELSE → use the default
Before you switch it on
Illustrative.
And when you need more
Controls that are already there.
Stop paying twice
Repeated questions are answered from a cache, and the saving is shown in dollars.
Open Cache →Keep sensitive data out
Mask or block personal data, secrets and injection attempts before a request leaves.
Open Guardrails →Check answers are still good
A judge scores a sample of answers, so a cheaper model is a measured choice.
Open Quality →Manage prompts safely
Version, roll back and A/B test prompts without redeploying.
Open Prompts →Give teams their own keys
Teams get limits of their own, and create their own keys, capped and expiring.
Open Teams →Get told, not surprised
Alerts on budgets and failures to email, Slack or a webhook.
Open Alerts →Send finance a statement
Usage by team, project or model, ready to export or schedule.
Open Reports →Prove who did what
A tamper-evident audit log you can verify and export.
Open Audit log →Get started
Five steps. About five minutes.
- Sign in.Use your CEEZ AI account. Sign in →
- Add a connection.Go to Connections, pick your provider (OpenAI, Anthropic, Gemini, a local model…), paste its key, choose the models you allow and press Test. Open Connections →
- Create a Compass key.Go to My keys, name the key and pick the project it belongs to. Copy it: it is shown once. Open My keys →
- Point your app at Compass.Change the base URL and use the Compass key. Nothing else changes.
- Watch it work.Your first call appears in the call log within seconds; Operations and Tokenomics fill in from there. Open the call log →
from openai import OpenAI client = OpenAI( base_url="https://compass.example.com/v1", # the change api_key="ck_…", # Compass key ) reply = client.chat.completions.create( model="openai/gpt-4o-mini", messages=[{"role": "user", "content": "Hello"}], )
New to Compass and no account yet? Request access. Any OpenAI-compatible SDK, framework or tool works, including streaming and tool calls.
Who uses it for what
Find your job, go straight to it.
Engineering and platform
Finance and operations
Security and compliance
Safe to run
Your data stays yours.
You decide whether prompts are stored at all. Stored secrets and backups are encrypted, outbound calls are contained, and in shared deployments each organisation is isolated from every other.
Will my existing code work?
If it talks to an OpenAI-compatible API, change the base URL and the key. Chat completions with streaming and tools, the Responses API, completions and embeddings are supported.
Does Compass store my prompts?
Your choice, per organisation, team or project: store text, store none, redact emails and keys in previews, keep calls for a set number of days. Tokens, cost and timing are always recorded; the words are optional.
What if a provider goes down?
Requests move to the fallback you set on errors, rate limits and outages, and a circuit breaker stops sending to a provider that keeps failing. The call log shows what happened.
What about a model whose price you do not know?
Set the price on the connection, per thousand or per million tokens, per model if you like. A model with no known price is shown as unpriced, never as free.
Who can create API keys?
Administrators and team leads, and, if you switch it on, any signed-in team member for their own team's projects, instantly or after approval. Keys are capped, expire by default and are shown once.
Ready
Open Compass and look at your traffic.
Sign in, add one connection, and your first call shows up live.