Skip to content

Guztia AI Infra Ops

Guztia /How it works

The approach, at a useful altitude

How It Works

What we actually build, described by function and the typical open-source technologies behind each layer. Also: where open-weight models genuinely still lose, and what you get to see before committing to anything.

01 The shape of it

Guztia reference architecture for enterprise AI Engineer tooling, product runtimes and CI pipelines all send requests to a gateway you control. The gateway routes them to open-weight models running on capacity in your own name, keeping a capped route to a premium vendor for the work that needs it. Every request is recorded underneath for cost attribution and audit. YOUR PEOPLE & PIPELINES Assistants & editors UNCHANGED CONFIGURATION Product runtimes YOUR APPLICATIONS CI & batch jobs AUTOMATED WORKLOADS CONTROL POINT — YOU OWN IT AI gateway ROUTING + FAILOVER SPEND LIMITS PER TEAM RATE LIMITS IDENTITY + ACCESS CACHING DATA-CLASS POLICY CHANGEABLE WITHOUT TOUCHING A LAPTOP Primary model · your capacity FIXED PRICE · REGION OF YOUR CHOICE Secondary model · your capacity FOR THE HARDER WORK Premium vendor · fallback CAPPED, MEASURED, OPTIONAL Cost & usage record — self-hosted EVERY REQUEST LOGGED · COST PER TEAM, PROJECT AND PERSON · LATENCY · QUALITY SCORES · RETAINED AUDIT TRAIL RUNS ON INFRASTRUCTURE IN YOUR NAME — AWS · ALIBABA CLOUD · GCP · AZURE · TENCENT · BARE METAL OPERATED BY GUZTIA UNDER A MONTHLY AGREEMENT — KEPT CURRENT · CAPACITY REBALANCED · INCIDENTS · MONTHLY COST REPORT
Fig. 01 — The gateway is the seam: left of it nothing changes, right of it everything can.

02 Four layers, all open source, all yours

You'll own every layer.

Typical technologies shown. Final selection depends on your workload, existing infrastructure and operational preferences — we propose the specific design under NDA before any commitment.

03 Why open source, specifically

Because the point is that you can fire us.

A managed service only its vendor can operate is not a solution to supplier lock-in. It is a smaller supplier with the same problem.

Everything we deploy is permissively licensed, runs in your accounts, and is documented publicly by people who are not us. That is not altruism — it is the only version of this business we think survives contact with a competent CTO, and it is why we charge for operating the arrangement rather than for access to it.

Not in the arrangement
Any Guztia-built software, any proprietary control layer, any margin on your usage, any hosted service of ours that your traffic must pass through.
You hold
The cloud accounts, the infrastructure code, the model weights, the usage records, the documentation.
We hold
A monthly invoice and the pager.
Leaving
Thirty days' notice. Nothing to extract, nothing to rebuild, no data to get back.

04 Where open weights still lose

We are not going to tell you the gap has closed.

It has narrowed a great deal. For the bulk of enterprise work — classification, extraction, summarisation, internal question answering, routine code — a good open-weight model is not distinguishable from a premium API in outcome, while costing a fraction. That is most of your volume, and it is the whole basis of the argument.

It is not all of it. There remain tasks where the premium models are simply better, and pretending otherwise would cost you more than you saved. That is exactly why every arrangement we build keeps a capped route to a premium vendor rather than burning the bridge.

Which of your workloads fall on which side is an empirical question about your work, not a matter of opinion — and answering it is the first thing an assessment does.

Long-horizon agents
Tasks running many steps with tool use and self-correction. Premium models hold together noticeably longer before drifting.
Genuinely novel reasoning
Unfamiliar problems rather than variations on well-covered ground. The gap is real and shows up clearly in testing.
Very long context
Open models advertise large windows; quality at the far end of them is a separate question. Test at the length you will actually use.
Low-resource languages
Coverage varies sharply. If you operate in one, it has to be tested specifically — a headline multilingual claim will not answer you.
Harder multimodal work
Open vision-language models are improving quickly but premium APIs remain ahead on difficult document and image reasoning.

THE RECOMMENDATION

Route by difficulty, not by loyalty

Bulk volume goes to capacity you control. Harder work goes to a larger model you also control. What neither can do well goes to a premium vendor, under a hard ceiling — and every one of those requests is recorded, so you can see precisely what the escape hatch costs and whether it is shrinking. In the arrangements we run, it shrinks.

See what your AI actually costs.

Thirty minutes, free, no deck. Bring whatever you know about what your teams send to model APIs — even a rough monthly number is enough. You leave with a view on what a gateway, routing and cost visibility would look like for your organisation.