Skip to content

Guztia AI Infra Ops

Guztia /Services /AI Infrastructure

SVC-01 · One-time engagement

AI Infrastructure Setup

Two to four weeks. At the end of it, every AI request routes through a gateway you control, every call is traced, cost is visible per team — and your engineers have changed one environment variable. Self-hosted models where the economics justify it; commercial APIs where they don't.

01 What gets built

Guztia reference architecture for enterprise AI Engineer tooling, product runtimes and CI pipelines all send requests to a gateway you control. The gateway routes them to open-weight models running on capacity in your own name, keeping a capped route to a premium vendor for the work that needs it. Every request is recorded underneath for cost attribution and audit. YOUR PEOPLE & PIPELINES Assistants & editors UNCHANGED CONFIGURATION Product runtimes YOUR APPLICATIONS CI & batch jobs AUTOMATED WORKLOADS CONTROL POINT — YOU OWN IT AI gateway ROUTING + FAILOVER SPEND LIMITS PER TEAM RATE LIMITS IDENTITY + ACCESS CACHING DATA-CLASS POLICY CHANGEABLE WITHOUT TOUCHING A LAPTOP Primary model · your capacity FIXED PRICE · REGION OF YOUR CHOICE Secondary model · your capacity FOR THE HARDER WORK Premium vendor · fallback CAPPED, MEASURED, OPTIONAL Cost & usage record — self-hosted EVERY REQUEST LOGGED · COST PER TEAM, PROJECT AND PERSON · LATENCY · QUALITY SCORES · RETAINED AUDIT TRAIL RUNS ON INFRASTRUCTURE IN YOUR NAME — AWS · ALIBABA CLOUD · GCP · AZURE · TENCENT · BARE METAL OPERATED BY GUZTIA UNDER A MONTHLY AGREEMENT — KEPT CURRENT · CAPACITY REBALANCED · INCIDENTS · MONTHLY COST REPORT
Fig. 01 — The gateway is the seam: left of it nothing changes, right of it everything can.

02 Components

All of it open source. None of it ours.

There is no Guztia software in this stack and no licence to renew. If you stop working with us, everything keeps running.

LayerWhat we deployWhy it matters
Serving vLLM · SGLang Turns rented GPU-hours into tokens at a cost you can forecast. Continuous batching is where the economics of self-hosting actually work.
Gateway LiteLLM Proxy One endpoint in front of every model — budgets, keys, fallbacks, routing. Makes everything else swappable.
Observability Langfuse (self-hosted) Traces, cost attribution, latency, eval scores. Self-hosted, so the record stays inside your company.
Orchestration Kubernetes · Docker Compose Matched to your scale. We will not sell you a control plane you do not need.
Compute Your cloud · reserved · spot · bare metal Placed where your data and users already are. Sized against measured traffic.
Identity Your existing corporate identity Cost attribution lands on a real person and team; revoking access revokes model access too.
Secrets Your existing secret store Keys rotate on a schedule. Whatever you already use, we use.
Everything Terraform · Helm · your CI In your repository, reviewed in your pull requests. We do not hold the state file.

03 How it runs

Four weeks, and week one is mostly measurement.

We do not start by deploying anything. We start by finding out what you actually send to model APIs today, because that number is usually a surprise.

Week 1 — Measure

What models are in use, by whom, for what, at what volume and what cost. Which calls are latency-sensitive and which are batch. What data classes are involved and what your legal position on them actually is.

Output: a baseline you can argue with, and a recommendation on which workloads move first.

Week 2 — Serve

GPU capacity provisioned, models deployed, endpoints benchmarked against your traffic shape rather than a synthetic one. Tuned until the numbers are honest.

Output: an inference endpoint with measured throughput, latency and cost per million tokens.

Week 3 — Route & trace

Gateway in front with routing rules, per-team keys, budgets, rate limits and a vendor fallback. Tracing behind it, wired to SSO so every trace carries a real identity. Dashboards built for the questions finance will ask.

Output: one endpoint, every call traced and costed.

Week 4 — Migrate & hand over

Engineers cut over — two environment variables, one short internal doc. CI and services follow. Then the runbook: how to add a model, roll one back, scale the pool, read the dashboards, and what to do at 3am when a node dies.

Output: the stack, the code, the runbook, and a team that can operate it without us.

04 The engineer-facing change

This is the whole migration.

Not a rewrite, not a new SDK, not a platform your team has to learn. Everything already speaks the same widely-adopted API format — the assistants, editors, frameworks, your own service code, your CI evals. The gateway answers in that format.

Which is also why the reverse is true: if you ever want to go back to a vendor API, that is the same two lines in the other direction. We think that matters. A migration you cannot reverse is not a migration, it is a different lock-in.

~/.zshrc — engineer workstation diff
- OPENAI_BASE_URL=https://api.openai.com/v1 - OPENAI_API_KEY=sk-proj-… # billed per token, untraced   + OPENAI_BASE_URL=https://llm.yourco.internal + OPENAI_API_KEY=sk-… # your gateway, your budget, your records
That is the entire migration on the engineer's side. Everything behind the URL — which model answers, which GPU it runs on, what it costs, who asked — is now yours to change.

05 What it costs

SCOPE A

Single model, single team

€8k–€12k

One open-weight model, gateway, tracing, one team migrated. Sized to your current infrastructure. Two to three weeks.

SCOPE B

Multi-model, org-wide

€15k–€25k

Several models across tiers, SSO, per-team budgets, autoscaling, vendor fallback, whole engineering org migrated. Three to four weeks.

SCOPE C

Regulated or cross-border

Quoted

Data residency constraints, audit requirements, mainland China presence, air-gapped environments. Scoped after week one.

Fixed fee, not time and materials. GPU and cloud costs are yours and billed by your provider directly — we never resell compute, because a margin on your inference would give us an interest in you using more of it. We model the break-even from your actual traffic before recommending anything — bring your invoice to the call.

06 The questions everyone asks

Do we need our own GPUs?
No. Most clients rent — reserved instances for the steady load, spot or serverless for bursts. Buying hardware only makes sense at sustained high utilisation, and we will tell you plainly when you are nowhere near that. Sometimes the honest answer after week one is "keep using the vendor API for now, but put the gateway in so you can see what you're spending."
Can we keep OpenAI or Anthropic?
Yes, and most clients do. The gateway routes bulk traffic to self-hosted models and sends the hard cases to a frontier model, with a budget cap on that path. You get the cost curve of self-hosting and the ceiling of the best available model, and you can finally see the split.
Which model should we run?
Whichever wins on your evals, which is not the same as whichever wins on a leaderboard. In practice a strong open-weight variant handles most of the volume and something larger handles the rest. We benchmark against your traffic in week two, and we expect the answer to change within a year — that is what the gateway is for.
What if the model is worse?
For some tasks it will be, and you should know which ones before you migrate them. That is the point of measuring first and of keeping a fallback route. Nothing moves to a self-hosted model until it passes your evals.
What happens if we stop working with you?
Everything keeps running. The infrastructure is in your cloud account, the code is in your repository, the tools are open source, and the runbook is written for someone who is not us. There is no Guztia component to remove.
Have you actually done this?
We run this exact arrangement in our own production and have operated infrastructure for other organisations since 2015. Ask on the call and we will walk you through a live instance — the real reporting and controls, not a slide.

See what your AI actually costs.

Thirty minutes, free, no deck. Bring whatever you know about what your teams send to model APIs — even a rough monthly number is enough. You leave with a view on what a gateway, routing and cost visibility would look like for your organisation.