Skip to content

Guztia AI Infra Ops

Guztia /Services /AI Infrastructure

SVC-01 · One-time engagement

AI Infrastructure Setup

Two to four weeks. At the end of it, open-weight models are serving on infrastructure you own, every request routes through a gateway you control, every call is traced, and your engineers have changed one environment variable.

01 What gets built

Guztia reference architecture for enterprise AI Engineer tooling, product runtimes and CI pipelines all send requests to a gateway you control. The gateway routes them to open-weight models running on capacity in your own name, keeping a capped route to a premium vendor for the work that needs it. Every request is recorded underneath for cost attribution and audit. YOUR PEOPLE & PIPELINES Assistants & editors UNCHANGED CONFIGURATION Product runtimes YOUR APPLICATIONS CI & batch jobs AUTOMATED WORKLOADS CONTROL POINT — YOU OWN IT AI gateway ROUTING + FAILOVER SPEND LIMITS PER TEAM RATE LIMITS IDENTITY + ACCESS CACHING DATA-CLASS POLICY CHANGEABLE WITHOUT TOUCHING A LAPTOP Primary model · your capacity FIXED PRICE · REGION OF YOUR CHOICE Secondary model · your capacity FOR THE HARDER WORK Premium vendor · fallback CAPPED, MEASURED, OPTIONAL Cost & usage record — self-hosted EVERY REQUEST LOGGED · COST PER TEAM, PROJECT AND PERSON · LATENCY · QUALITY SCORES · RETAINED AUDIT TRAIL RUNS ON INFRASTRUCTURE IN YOUR NAME — AWS · ALIBABA CLOUD · GCP · AZURE · TENCENT · BARE METAL OPERATED BY GUZTIA UNDER A MONTHLY AGREEMENT — KEPT CURRENT · CAPACITY REBALANCED · INCIDENTS · MONTHLY COST REPORT
Fig. 01 — The gateway is the seam: left of it nothing changes, right of it everything can.

02 Components

All of it open source. None of it ours.

There is no Guztia software in this stack and no licence to renew. If you stop working with us, everything keeps running.

LayerWhat we deployWhy it matters
Serving Inference engine on your GPUs Turns rented GPU-hours into tokens at a cost you can forecast.
Gateway AI gateway (open source) One endpoint in front of every model — budgets, keys, fallbacks, routing. Makes everything else swappable.
Observability Usage & cost record (open source, self-hosted) Traces, cost attribution, latency, eval scores. Self-hosted, so the record stays inside your company.
Orchestration Matched to your scale Matched to your scale. We will not sell you a control plane you do not need.
Compute Your cloud · reserved · spot · bare metal Placed where your data and users already are. Sized against measured traffic.
Identity Your existing corporate identity Cost attribution lands on a real person and team; revoking access revokes model access too.
Secrets Your existing secret store Keys rotate on a schedule. Whatever you already use, we use.
Everything Infrastructure as code you keep In your repository, reviewed in your pull requests. We do not hold the state file.

03 How it runs

Four weeks, and week one is mostly measurement.

We do not start by deploying anything. We start by finding out what you actually send to model APIs today, because that number is usually a surprise.

Week 1 — Measure

What models are in use, by whom, for what, at what volume and what cost. Which calls are latency-sensitive and which are batch. What data classes are involved and what your legal position on them actually is.

Output: a baseline you can argue with, and a recommendation on which workloads move first.

Week 2 — Serve

GPU capacity provisioned, models deployed, endpoints benchmarked against your traffic shape rather than a synthetic one. Tuned until the numbers are honest.

Output: an inference endpoint with measured throughput, latency and cost per million tokens.

Week 3 — Route & trace

Gateway in front with routing rules, per-team keys, budgets, rate limits and a vendor fallback. Tracing behind it, wired to SSO so every trace carries a real identity. Dashboards built for the questions finance will ask.

Output: one endpoint, every call traced and costed.

Week 4 — Migrate & hand over

Engineers cut over — two environment variables, one short internal doc. CI and services follow. Then the runbook: how to add a model, roll one back, scale the pool, read the dashboards, and what to do at 3am when a node dies.

Output: the stack, the code, the runbook, and a team that can operate it without us.

04 The engineer-facing change

This is the whole migration.

Not a rewrite, not a new SDK, not a platform your team has to learn. Everything already speaks the same widely-adopted API format — the assistants, editors, frameworks, your own service code, your CI evals. The gateway answers in that format.

Which is also why the reverse is true: if you ever want to go back to a vendor API, that is the same two lines in the other direction. We think that matters. A migration you cannot reverse is not a migration, it is a different lock-in.

~/.zshrc — engineer workstation diff
- OPENAI_BASE_URL=https://api.openai.com/v1 - OPENAI_API_KEY=sk-proj-… # billed per token, untraced   + OPENAI_BASE_URL=https://llm.yourco.internal + OPENAI_API_KEY=sk-… # your gateway, your budget, your records
That is the entire migration on the engineer's side. Everything behind the URL — which model answers, which GPU it runs on, what it costs, who asked — is now yours to change.

05 What it costs

SCOPE A

Single model, single team

€8k–€12k

One open-weight model, gateway, tracing, one team migrated. Sized to your current infrastructure. Two to three weeks.

SCOPE B

Multi-model, org-wide

€15k–€25k

Several models across tiers, SSO, per-team budgets, autoscaling, vendor fallback, whole engineering org migrated. Three to four weeks.

SCOPE C

Regulated or cross-border

Quoted

Data residency constraints, audit requirements, mainland China presence, air-gapped environments. Scoped after week one.

Fixed fee, not time and materials. GPU and cloud costs are yours and billed by your provider directly — we never resell compute, because a margin on your inference would give us an interest in you using more of it. Teams spending €10k+/month on API calls typically see the build pay back within a quarter; bring your invoice to the assessment and we will say whether that is true for you.

06 The questions everyone asks

Do we need our own GPUs?
No. Most clients rent — reserved instances for the steady load, spot or serverless for bursts. Buying hardware only makes sense at sustained high utilisation, and we will tell you plainly when you are nowhere near that. Sometimes the honest answer after week one is "keep using the vendor API for now, but put the gateway in so you can see what you're spending."
Can we keep OpenAI or Anthropic?
Yes, and most clients do. The gateway routes bulk traffic to self-hosted models and sends the hard cases to a frontier model, with a budget cap on that path. You get the cost curve of self-hosting and the ceiling of the best available model, and you can finally see the split.
Which model should we run?
Whichever wins on your evals, which is not the same as whichever wins on a leaderboard. In practice a strong open-weight variant handles most of the volume and something larger handles the rest. We benchmark against your traffic in week two, and we expect the answer to change within a year — that is what the gateway is for.
What if the model is worse?
For some tasks it will be, and you should know which ones before you migrate them. That is the point of measuring first and of keeping a fallback route. Nothing moves to a self-hosted model until it passes your evals.
What happens if we stop working with you?
Everything keeps running. The infrastructure is in your cloud account, the code is in your repository, the tools are open source, and the runbook is written for someone who is not us. There is no Guztia component to remove.
Have you actually done this?
We run this exact arrangement in our own production and have operated infrastructure for other organisations since 2015. Ask on the call and we will walk you through a live instance — the real reporting and controls, not a slide.

Own the stack before the bill owns you.

Thirty minutes. Bring your current AI invoice. We look at what your teams send to model APIs, what it costs today, and what a self-hosted stack would cost instead. No deck, no discovery phase.