Skip to content

Guztia AI Infra Ops

AI Infrastructure Operations

Your AI bill is about to double.
We make it predictable.

We put self-hosted models on your cloud, behind a gateway you control, with cost visibility per team — then operate it monthly. Your engineers change one environment variable. Your finance team finally gets a forecast they can use.

~/.zshrc — engineer workstation diff
- OPENAI_BASE_URL=https://api.openai.com/v1 - OPENAI_API_KEY=sk-proj-… # billed per token, untraced   + OPENAI_BASE_URL=https://llm.yourco.internal + OPENAI_API_KEY=sk-… # your gateway, your budget, your records
That is the entire migration on the engineer's side. Everything behind the URL — which model answers, which GPU it runs on, what it costs, who asked — is now yours to change.

01 The problem

Someone else is paying for your inference. That ends.

Frontier labs have been selling tokens and seats below what the GPUs cost them, funded by capital rather than revenue. That gap closes eventually — through price rises, tighter limits, or both. The teams that built on a subsidy and never measured it will find out the hard way.

Cost

Nobody in the building can say what an agent run costs, which team spent the most last month, or what would happen to the bill if list prices doubled. The invoice arrives as one number.

Control

Every laptop, CI runner and service holds its own vendor key. There is no budget ceiling, no rate limit you set, no record of which prompts left the building or what came back.

Lock-in

Switching models means a change-management project across every machine and pipeline. So you don't switch — and the vendor knows it. That is the whole business model.

Residency

Legal asks where the prompts go. Nobody has an answer that survives an audit, and "the vendor says they don't train on it" is not one.

02 What we build

One change for your engineers. Full control for you.

Four open-source components, wired correctly, on infrastructure you own — then kept alive by people who have done it in production.

Guztia reference architecture for enterprise AI Engineer tooling, product runtimes and CI pipelines all send requests to a gateway you control. The gateway routes them to open-weight models running on capacity in your own name, keeping a capped route to a premium vendor for the work that needs it. Every request is recorded underneath for cost attribution and audit. YOUR PEOPLE & PIPELINES Assistants & editors UNCHANGED CONFIGURATION Product runtimes YOUR APPLICATIONS CI & batch jobs AUTOMATED WORKLOADS CONTROL POINT — YOU OWN IT AI gateway ROUTING + FAILOVER SPEND LIMITS PER TEAM RATE LIMITS IDENTITY + ACCESS CACHING DATA-CLASS POLICY CHANGEABLE WITHOUT TOUCHING A LAPTOP Primary model · your capacity FIXED PRICE · REGION OF YOUR CHOICE Secondary model · your capacity FOR THE HARDER WORK Premium vendor · fallback CAPPED, MEASURED, OPTIONAL Cost & usage record — self-hosted EVERY REQUEST LOGGED · COST PER TEAM, PROJECT AND PERSON · LATENCY · QUALITY SCORES · RETAINED AUDIT TRAIL RUNS ON INFRASTRUCTURE IN YOUR NAME — AWS · ALIBABA CLOUD · GCP · AZURE · TENCENT · BARE METAL OPERATED BY GUZTIA UNDER A MONTHLY AGREEMENT — KEPT CURRENT · CAPACITY REBALANCED · INCIDENTS · MONTHLY COST REPORT
Fig. 01 — The gateway is the seam: left of it nothing changes, right of it everything can.

A · ZERO DISRUPTION

Zero-disruption migration

Your AI assistants, your editors, your frameworks, your own SDK calls, your CI evals — they already speak one API shape. The gateway answers in that shape, so the migration is a base URL and a key, not a rewrite.

B · MODEL FREEDOM

Switch models without a project

A new open-weight release ships, benchmarks well on your evals, and goes live in a routing rule that afternoon. Nobody reinstalls anything. Cheap models take the bulk traffic, expensive ones take what needs them.

C · COST VISIBILITY

Cost per team, per call

Every request is traced: who asked, which model answered, how long, how much, how good. Cost per team stops being a guess and becomes a line in a dashboard you can hand to finance.

03 What you get back

This is what cost visibility looks like.

Not a monthly invoice. A span waterfall, per call, with the model, the tokens, the tool hops and the cost attached — kept as long as your retention policy says.

How the record works

05 Why us

We ran this stack on ourselves first.

Guztia has been doing cloud operations since 2015 — migrations, Kubernetes, Terraform, 24/7 retainers, mostly across APAC and China. The discipline transfers directly. A GPU node that OOMs at 3am is not a new category of problem.

Who is behind this

2015
Operating clouds since
Alibaba Cloud MVP, consecutive
100+
Organisations served
129
Technical posts, no ghostwriter
  • We run this exact arrangement in our own production — it is what we operate daily, not a demonstration built for the pitch.
  • Alibaba Cloud MVP eight years running, which is why deployments in Asia-Pacific are routine here rather than a research project.
  • China cross-border since 2017: if your models need to answer inside the mainland, that is an old problem here, not a new one.
  • Two named engineers — Alberto Roura and Xavier Macia — plus a specialist network. You always know who is on the incident.

06 Where it runs

Your account, your region, your commitment discounts.

We deploy into infrastructure you already pay for. No Guztia-hosted middle layer, no per-token margin, nothing to migrate off if you fire us.

07 Still here

The China practice didn't go anywhere.

Business VPN, site acceleration, ICP filing, cross-border connectivity, Salesforce access from the mainland. Same team, same lines, same licences — now with the added question of where your models are allowed to run and whose jurisdiction the prompts land in.