Skip to content

Guztia AI Infra Ops

AI Infrastructure Operations

Control the infrastructure
behind your AI.

One gateway for every model. Predictable cost, traceable usage, data residency and failover — deployed in your cloud and operated by Guztia.

Commercial APIs + self-hosted models · EU + APAC + China · Open source · Your infrastructure · No lock-in

~/.zshrc — engineer workstation diff
- OPENAI_BASE_URL=https://api.openai.com/v1 - OPENAI_API_KEY=sk-proj-… # billed per token, untraced   + OPENAI_BASE_URL=https://llm.yourco.internal + OPENAI_API_KEY=sk-… # your gateway, your budget, your records
That is the entire migration on the engineer's side. Everything behind the URL — which model answers, which GPU it runs on, what it costs, who asked — is now yours to change.

01 The problem

AI usage is growing faster than companies can control it.

Multiple models, multiple vendors, every team holding its own key. No budget ceiling, no per-team breakdown, no record of where prompts go. The invoice arrives as one growing number.

Cost

Nobody can say what an agent run costs, which team spent the most, or what the bill looks like next quarter.

Control

No central gateway, no budget ceiling, no rate limit you set. Every laptop and CI runner holds a vendor key.

Lock-in

Switching models means a change programme across every service. So you don't switch — and the vendor knows it.

Residency

Legal asks where the prompts go. Nobody has an answer that survives an audit.

02 What we build

One change for your engineers. Full control for you.

Open-source components — LiteLLM, Langfuse, vLLM — wired correctly on infrastructure you own. Typical technologies; final selection depends on workload.

Guztia reference architecture for enterprise AI Engineer tooling, product runtimes and CI pipelines all send requests to a gateway you control. The gateway routes them to open-weight models running on capacity in your own name, keeping a capped route to a premium vendor for the work that needs it. Every request is recorded underneath for cost attribution and audit. YOUR PEOPLE & PIPELINES Assistants & editors UNCHANGED CONFIGURATION Product runtimes YOUR APPLICATIONS CI & batch jobs AUTOMATED WORKLOADS CONTROL POINT — YOU OWN IT AI gateway ROUTING + FAILOVER SPEND LIMITS PER TEAM RATE LIMITS IDENTITY + ACCESS CACHING DATA-CLASS POLICY CHANGEABLE WITHOUT TOUCHING A LAPTOP Primary model · your capacity FIXED PRICE · REGION OF YOUR CHOICE Secondary model · your capacity FOR THE HARDER WORK Premium vendor · fallback CAPPED, MEASURED, OPTIONAL Cost & usage record — self-hosted EVERY REQUEST LOGGED · COST PER TEAM, PROJECT AND PERSON · LATENCY · QUALITY SCORES · RETAINED AUDIT TRAIL RUNS ON INFRASTRUCTURE IN YOUR NAME — AWS · ALIBABA CLOUD · GCP · AZURE · TENCENT · BARE METAL OPERATED BY GUZTIA UNDER A MONTHLY AGREEMENT — KEPT CURRENT · CAPACITY REBALANCED · INCIDENTS · MONTHLY COST REPORT
Fig. 01 — The gateway is the seam: left of it nothing changes, right of it everything can.

A · ZERO DISRUPTION

Same API, different economics

The gateway answers in the same API shape your tools already speak. Migration is a base URL and a key, not a rewrite.

B · MODEL FREEDOM

Switch models in a routing rule

New release benchmarks well on your evals — it goes live that afternoon. Cheap models take bulk traffic, expensive ones take what needs them.

C · COST VISIBILITY

Cost per team, per call

Every request traced: who asked, which model, how long, how much. Cost per team becomes a dashboard line you hand to finance.

03 What you get back

Cost visibility, per call.

A span waterfall — model, tokens, tool hops, cost — kept as long as your retention policy says. Powered by Langfuse, self-hosted.

How the stack works

05 Why us

Cloud ops since 2015. AI infrastructure since 2023.

Migrations, Kubernetes, Terraform, 24/7 retainers across APAC and China. The discipline — measure, control, operate, report — is the same. The workload is new.

Alberto Roura & Xavier Macia

2015
Cloud operations since
2023
AI infrastructure practice
100+
Infrastructure customers
Alibaba Cloud MVP
  • We run this arrangement in our own production — not a demo built for the pitch.
  • China cross-border since 2017 — if your models need to answer inside the mainland, that is an old problem here.
  • Two partners work on your engagement directly. You always know who is on the incident.

06 Where it runs

Your account, your region, your commitment discounts.

We deploy into infrastructure you already pay for. No Guztia-hosted middle layer, no per-token margin, nothing to migrate off if you fire us.

07 AI infrastructure across jurisdictions

EU, APAC and mainland China — one control plane.

Few European AI infrastructure firms can operate across jurisdictional boundaries. We have been doing exactly that since 2017 — enterprise VPN, site acceleration, ICP filing, cross-border connectivity — and now the payload is AI.

EU workloads route to EU models and infrastructure. China workloads route to approved local providers. Sensitive workloads stay on private inference. The application sees one API; policy decides where the prompt lands.

EU
OpenAI · Anthropic · Mistral · EU-hosted open-weight
China
Qwen · DeepSeek · local model infrastructure
Sensitive
Private inference · air-gapped · your hardware
Since
2017 · AWS China · Alibaba · Tencent · Huawei