AI monitoring · cost tracking · model routing

See every AI call. Control what it costs. Pick the right model.

CEEZ Compass sits between your apps and AI providers like OpenAI and Anthropic. Sign in to watch live traffic, see every dollar by team and model, and decide which model answers which request.

Works with any OpenAI-compatible app · 38 ready-made providers · No account? Request access

Monitoring

Know the moment AI misbehaves.

Every request through Compass is recorded: how long it took, whether it worked, which provider answered and what it cost. You get a live picture instead of hearing about it from a customer.

  • Live call log: each request with its status, tokens, cost and the reason it went where it did
  • Health and latency: success rate, typical and slow answers, time to first word, by model
  • Reliability: failovers, retries and rate limits, and which provider is struggling
  • Alerts: email, Slack or webhook when errors or budgets cross a line

The benefit: a provider outage becomes a line in a log, because Compass has already moved traffic to your backup.

Open Operations →   Open the call log →

Tokenomics

Know what AI costs, and who spent it.

AI is billed in tokens. Compass counts them on every call and turns them into dollars by team, app, key and model, using the prices you set, including the cheaper rate for cached input.

  • Spend and usage over time, with a forecast to month end
  • Models and callers: which model and which app cost the most
  • Budgets and limits by day or month, per key, project and team: refuse at the limit, or switch to a cheaper model
  • Savings advice from your own traffic: repeated questions, over-qualified models
  • Chargeback reports per team, ready for finance

The benefit: one provider bill becomes a statement per team, and the next saving is a suggestion, not a hunt.

Open Tokenomics →   Set a budget →

Model routing

Send each request to the right model.

You rarely need the most expensive model for every question. Write simple rules once, and Compass applies them to every request, with no change to your apps.

  • By what the request is: short or long, simple or complex, which tools it uses
  • By what is happening: a provider failing, slow, or a team's budget used up
  • By policy: keep certain data in certain regions, or on certain providers
  • Cheap first, escalate if needed, or split traffic to compare two models
  • Test before you commit: replay last week's traffic, or run a rule in shadow mode

The benefit: pay for the strong model only when a request needs it, and never be stuck on one provider.

Open Model routing →   Add a connection →

Get started

Five steps. About five minutes.

  1. Sign in.Use your CEEZ AI account. Sign in →
  2. Add a connection.Go to Connections, pick your provider (OpenAI, Anthropic, Gemini, a local model…), paste its key, choose the models you allow and press Test. Open Connections →
  3. Create a Compass key.Go to My keys, name the key and pick the project it belongs to. Copy it: it is shown once. Open My keys →
  4. Point your app at Compass.Change the base URL and use the Compass key. Nothing else changes.
  5. Watch it work.Your first call appears in the call log within seconds; Operations and Tokenomics fill in from there. Open the call log →
Step 4: your app
from openai import OpenAI

client = OpenAI(
    base_url="https://compass.example.com/v1",  # the change
    api_key="ck_…",                               # Compass key
)
reply = client.chat.completions.create(
    model="openai/gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
)

New to Compass and no account yet? Request access. Any OpenAI-compatible SDK, framework or tool works, including streaming and tool calls.

Safe to run

Your data stays yours.

You decide whether prompts are stored at all. Stored secrets and backups are encrypted, outbound calls are contained, and in shared deployments each organisation is isolated from every other.

Read the security overview →

Will my existing code work?

If it talks to an OpenAI-compatible API, change the base URL and the key. Chat completions with streaming and tools, the Responses API, completions and embeddings are supported.

Does Compass store my prompts?

Your choice, per organisation, team or project: store text, store none, redact emails and keys in previews, keep calls for a set number of days. Tokens, cost and timing are always recorded; the words are optional.

What if a provider goes down?

Requests move to the fallback you set on errors, rate limits and outages, and a circuit breaker stops sending to a provider that keeps failing. The call log shows what happened.

What about a model whose price you do not know?

Set the price on the connection, per thousand or per million tokens, per model if you like. A model with no known price is shown as unpriced, never as free.

Who can create API keys?

Administrators and team leads, and, if you switch it on, any signed-in team member for their own team's projects, instantly or after approval. Keys are capped, expire by default and are shown once.

Ready

Open Compass and look at your traffic.

Sign in, add one connection, and your first call shows up live.