For the CTO
Sovereign intelligence for the CTO
Deploy models, agents and workflows on infrastructure you control.
The CTO decides what the AI stack is made of and how much of it the company can change later. The technical risks are specific: inference you cannot move, retrieval you cannot port, agents with ambient credentials and observability that stops at the vendor's edge. This page covers the architecture decisions and the questions that expose lock-in.
What a CTO worries about
Inference economics and placement
Serving models well is a systems problem: memory for the KV cache, batching, queueing and GPU supply. Choosing hosted, dedicated or self-run inference changes cost, latency and who can see the traffic. See the vLLM write-up.
A control plane welded to one cloud
If routing, policy and observability only work inside one provider, then moving workloads means rebuilding the control plane. Separate the data plane from the control plane early.
Agents with too much authority
Tool calls made with a shared service account are an incident waiting to happen. Each agent needs its own identity, scoped credentials and a policy check on every call.
Evaluation debt
Without your own test sets, model choice is a guess and every upgrade is a risk. Teams that skip evaluation cannot tell whether a change made things better.
Protocol and format churn
Agent protocols and tracing conventions are still moving. Betting on a proprietary format today creates tomorrow's migration project.
What a Sovereign Intelligence Platform gives you
A layered architecture you can reason about
Six layers from infrastructure to solutions, with governance running through all of them, give a common map for design reviews. See the reference architecture.
Deployment flexibility
The platform is designed for cloud, private, on-premise and hybrid deployment, with dedicated cloud and GPU references for the infrastructure layer.
One API across many models
Connect is an entry point for a single API to 50+ model providers with routing, failover and cost tracking.
Policy enforced at the call
Runtime governance is designed so the policy decision happens on each tool call and model request, not in a review afterwards. See runtime governance.
Agents and workflows as governed objects
Studio and Nexus are entry points for building agents and workflows and for capturing and enforcing policy on agent actions.
Capabilities are described as what the platform is designed to let you do. For what is true today and what is not claimed, see the trust centre.
Questions to ask any vendor
Use this as a checklist in any evaluation, ours included. Each question comes with what a good answer looks like.
01Which parts of the stack can we run ourselves, and which are only available as the vendor's hosted service?
A good answer: A component-by-component table of self-hostable and hosted-only parts, with licence terms for the self-hosted ones.
02How is inference served: engine, batching, quantisation options and GPU types?
A good answer: Named engines and configuration options, with the freedom to bring our own serving stack. Be cautious of throughput claims without a stated workload.
03Can we bring our own model weights and fine-tunes, and are they ever copied to your systems?
A good answer: Yes, with a documented import path and a statement of where weights are stored, who can access them and how they are deleted.
04How do agents authenticate to tools, and where are credentials stored?
A good answer: Per-agent identities with short-lived, scoped credentials held in a secrets manager. Shared long-lived keys in prompts or config are a fail.
05Which agent and tool protocols do you support, such as the Model Context Protocol?
A good answer: Support for open protocols with a stated version. MCP now sits under the Linux Foundation (opens in a new tab), which reduces proprietary risk, but ask how tool permissions are enforced.
06What telemetry do you emit, and in what format?
A good answer: Traces for model calls, tool calls and agent steps in an open format such as OpenTelemetry, with a note that the GenAI conventions are still in Development status.
07How do we run our own evaluations and gate releases on them?
A good answer: An evaluation harness or API that accepts our datasets and metrics, and release gates that use the results.
08What are the failure modes of your policy engine, and does it fail closed?
A good answer: A documented behaviour for engine outage, timeout and malformed policy: sensitive actions are denied by default, with an alert.
09How do you isolate tenants and workloads, including GPU and cache layers?
A good answer: A description of network, compute, storage and cache isolation, with dedicated options and an explanation of what is shared in multi-tenant mode.
10How do we reproduce a model response for debugging or an incident review?
A good answer: Stored request parameters, model version, prompt template version and retrieved context for each call, so a result can be replayed, with a note on where sampling makes exact reproduction impossible.
11What are the rate limits, latency targets and regional availability, in writing?
A good answer: Stated limits and regions, with any targets in the contract rather than a marketing page. If a number is not available, expect a placeholder, not an invented figure.
12What do we export if we leave, and in which formats?
A good answer: Workflows, prompts, evaluations, vector data, logs and configuration in documented formats, with an export API rather than a support ticket.
Recommended reading
- Guide: choose infrastructure
- Guide: select and evaluate models
- Infrastructure sovereignty
- Sovereign AI reference architecture (blog)
- Platform: infrastructure layer
Not sure where to start? Take the readiness assessment or read the build guide.
Frequently asked questions
Should we self-host models or use hosted APIs?
It depends on data classification, volume and team capacity. Many organisations run both, routing sensitive workloads to infrastructure they control and everything else to hosted providers under policy.
Is a vector database enough for context?
Usually not alone. Good context combines retrieval, access control, lineage and sometimes a knowledge graph and memory. See the guide step on data and context.
How do we avoid lock-in with agents?
Use open protocols for tools, keep prompts, evaluations and workflows exportable, and keep identity and policy in a layer you can move.
Build with control: for the CTO
Start with one entry point. Add intelligence, agents, workflows and infrastructure as you prove value. Or read the step-by-step build guide and take the readiness assessment.