Context & tools
The material the work starts from: the sources, documents and tools a team already relies on.
TL;DR: An application-level gateway (ALG) understands the protocol it serves; not just IP packets. In 2026, the most-deployed ALG category is the AI gateway: a specialised proxy for LLM traffic that applies routing, caching, eval, audit, and cost-control policy at the chat-completion protocol layer.
| Category | What it does | Examples |
|---|---|---|
| Application-level gateway (ALG) | A proxy that understands the application protocol it serves. HTTP, SIP, FTP, SMTP, and applies inspection, transformation, and policy at that layer. | Traditional ALGs in firewalls: SIP ALG, FTP ALG. Cloud-era ALGs include Cloudflare Workers, AWS API Gateway, Kong, NGINX Plus. |
| AI gateway | A specialised ALG for LLM traffic; speaks the chat-completion / embeddings / tool-use protocol, applies routing, caching, eval, and cost policy. | Swfte, OpenRouter, Portkey, LiteLLM, TrueFoundry, Cloudflare AI Gateway. |
| MCP gateway | A specialised ALG for Model Context Protocol, and speaks the MCP wire protocol, applies tool-call auth, audit, and rate limit. | Swfte (consolidated), Anthropic mcp-proxy, Cloudflare AI Gateway. |
Route requests across Anthropic, OpenAI, Google, DeepSeek, Grok based on cost, latency, quality, or compliance policy. The ALG-style request transformation handles per-provider auth, retries, and observability.
A gateway sits in the request path and can normalise + cache prompt prefixes regardless of upstream provider, layering on top of provider-native caching for compound savings.
Track and enforce monthly budgets per team, per project, per user. Hard cut-off, soft alert, or routing-to-cheaper-tier on threshold approach.
Mirror traffic to a second model for offline comparison. Promote based on eval pass-rate. Catch regressions before they hit production.
Every request logged with caller identity, prompt content (or redacted), response, latency, cost, and model. Exports to SIEM. Required for SOC2, HIPAA, EU AI Act conformity.
Pre-prompt and post-response inspection. strip PII, block disallowed content, enforce response schema. Applied uniformly regardless of which provider is upstream.
An application-level gateway (ALG) is a proxy that understands the application protocol it serves, HTTP, SIP, FTP, SMTP, MQTT, gRPC: rather than operating only at L4 (TCP/UDP). The ALG can inspect payloads, transform requests, enforce policy at the protocol level, and provide telemetry that lower-level proxies cannot.
An AI gateway is a specialised ALG for the chat completion / embeddings / tool-use / MCP wire protocols used by LLMs. It speaks the upstream provider's protocol natively; OpenAI-format, Anthropic Messages format, Gemini's GenerateContent, and and applies policy at that layer. Routing decisions, caching, prompt transformations, response schema enforcement, and cost attribution all happen at the protocol level.
Yes. they serve different layers. A generic API gateway (Kong, AWS API Gateway, Cloudflare) handles HTTP-level concerns: rate limiting, auth, TLS termination. An AI gateway adds LLM-specific concerns: provider routing, prompt caching, model fallback, token-cost attribution, eval, guardrails. Most production AI deployments run both, the generic API gateway in front of the AI gateway.
Yes. Swfte is an AI-specialised application-level gateway plus an agent runtime. The gateway speaks chat-completion, embeddings, tool-use, and MCP protocols natively, applies routing / caching / eval / cost-control policy at the protocol level, and exposes a single OpenAI-compatible HTTP API to applications.
A reverse proxy (NGINX, HAProxy, Envoy) routes traffic at L4 / L7 based on hostname, path, headers. It does not understand the AI-specific payload: it cannot decide "this prompt looks like a code-gen task, route to Claude Opus" or "cache this prefix at the protocol level". An AI gateway adds that protocol-aware layer.
For OSS / self-hosted: LiteLLM. For managed / pay-as-you-go: Swfte free tier and OpenRouter both cost nothing up front and bill on usage. For enterprise: Portkey, TrueFoundry, and Swfte enterprise tier are competitive at scale.
Yes; Cloudflare AI Gateway, generally available since 2024 and continuously evolved through 2025-26. Strong fit for teams already running on Cloudflare Workers for edge compute. Less mature on agent runtime, eval harness, and per-team budget enforcement compared to dedicated AI platforms.
Because the lessons from a decade of API gateways apply directly. The same primitives that mattered for HTTP, and routing, caching, auth, rate limit, audit, observability. matter for AI. The infrastructure community is converging on calling AI gateways what they are: application-level gateways specialised for AI traffic. Adopting the term clarifies the architecture and accelerates the procurement conversation with infrastructure teams.
Swfte speaks chat-completion, embeddings, tool-use, and MCP natively. Routing, caching, audit, cost control at the protocol layer.
Free tier · OpenAI-compatible API · On-prem available
Connection map / Application Gateway
Inspect routing, data handling and budget boundaries without sending a request. Illustrative authored example · snapshot · 7 September 2026. Supplied catalogue copy, not verified performance evidence; this preview does not execute a workflow.
Choose the team’s configured model route for a routine summarisation task.
Route record: approved provider and project attribution.
Apply the configured redaction rule before forwarding upstream.
Transformed input: customer contact becomes [EMAIL].
Use the chosen ceiling behaviour rather than silently continuing.
Policy example: stop the request and report the project budget boundary.
A spatial explanation of the workflow.
Select a stage or scroll to explore.
The material the work starts from: the sources, documents and tools a team already relies on.
The reasoning step. A model reads the context that was gathered and proposes something a person can act on.
The limits set before the work runs — what is in scope, what is refused, and who is asked when it is unclear.
The sequence itself: the order of steps, the handoffs between them, and where a person is required.
How the result reaches the people who use it, and the conditions under which it is allowed to run.
What is kept afterwards so a reviewer can retrace the decision: inputs, choices, and the person accountable.
CONCEPTUAL WORKFLOW MODEL — NOT A DEPICTION OF SYSTEM ARCHITECTURE
Illustrative authored example · snapshot · 7 September 2026. Capabilities and figures require confirmation.