Platform
Control Plane

One place to run every model, prompt, tool and dollar.

Model routing, prompt management, tool and MCP governance, usage and cost, evaluation, guardrails and audit — the whole operating layer for AI in one place, rather than seven tools and a spreadsheet.

One gatewayIn-region inferenceVersioned prompts
Why It Matters

Why It Matters

What changes when you run this at your size.

The sprawl is the problem, not the models

Most organisations do not have an AI problem, they have eleven of them: a different provider per team, prompts pasted into code, keys in environment variables nobody can inventory, and no consolidated view of what any of it costs. A control plane is what turns that into one estate you can actually operate, which is why this layer usually decides whether a programme survives its second year.

Changing model or provider should be a setting, not a project

Model pricing and capability move every few months. If your provider choice is compiled into applications, you cannot act on that, and you end up paying last year’s prices for last year’s quality. Routing through one gateway means switching provider, pinning a version or moving cheap work to a cheaper model is a configuration change rather than an engineering programme.

Cost per completed process, not cost per token

Token spend is not a number a finance function can act on. Attributing spend to the team, the model and the process it belongs to is, because it lets you compare what a process costs to run against the work it replaced. That comparison is what makes the second round of investment defensible.

What It Does

What It Does

The detail behind the claim.

Every request goes through one gateway rather than each application holding its own provider keys. Route work by cost, capability, context size or latency, so a straightforward classification does not run on your most expensive model. Pin an exact model version where you need stability, and fail over automatically to the next provider in your order when one errors or throttles. Your provider credentials sit in one place, and switching provider is a setting.

On the managed Bixie cloud, inference runs on our own models hosted in the region you operate in. By default no customer data is sent to an external model API — prompts, context and outputs are processed in-region on first-party models, so they stay inside the jurisdiction you chose. Commercial and open-weight models can still be added to the gateway when you want them, under the same routing and policy controls.

Instructions live in a versioned registry rather than in application code. Change one and it takes effect without a deployment, compare it against the version before it, and roll back to any earlier version when a change makes things worse. Promote a change through staging before it reaches production, and require review on the prompts that carry real consequence, so the text that governs an agent’s behaviour gets the same treatment as the code around it.

One registry of every connected system and every action it exposes, with its schema, its owner and its purpose recorded. Each action is graded by consequence — reading something, changing something, or destroying something — and every agent gets an explicit grant rather than blanket access. Credentials bind to the tool and the tenant, so agents act without ever holding your keys. Connections report their own health, and you find out that access lapsed before the person relying on it does.

Spend attributed to the organisation, the team, the person, the model and the process, so showback and chargeback come out of the platform rather than a reconciliation exercise. Set quotas and budgets per team or per application, get told when consumption starts diverging from its own pattern, and see a forecast of where the month is heading while there is still time to act on it.

Changes to a prompt or a model get tested before they reach anyone. Keep a set of representative cases, run a candidate change against them, and compare the results to the version in production. Sample live traffic for quality scoring, route anything doubtful to a person, and hold releases that regress against the cases you already agreed matter.

Policy applied in the path of the request rather than hoped for afterwards. Detect and mask sensitive data before it reaches a model, screen inbound content for instructions trying to redirect an agent, apply your content standards, and require a named approver on anything consequential. Blocked and held requests are recorded with the reason, so a pattern of near-misses is visible rather than invisible.

A complete record of who asked for what, which model answered, which tools were called and who approved them, filterable and exportable for whoever asks. Access is governed by your existing identity provider with role-based permissions, and retention is set to match your own obligations rather than a vendor default.

One gateway
Every model behind a single entry point
In-region inference
First-party models, no external API by default
Versioned prompts
Change, compare and roll back without a deploy
Graded by consequence
Read, change, destroy — per tool and per action
Spend by team and process
Attributed, not apportioned
By Function
More Platform

See it against your own process

Bring something that currently costs your team a day a month. We will build it on the call rather than describe it, and you can watch every step it takes.

Cloud, self-hosted, or your own model