AI Infrastructure Operations
Control the infrastructure
behind your AI.
One gateway for every model. Predictable cost, traceable usage, data residency and failover — deployed in your cloud and operated by Guztia.
Commercial APIs + self-hosted models · EU + APAC + China · Open source · Your infrastructure · No lock-in
- OPENAI_BASE_URL=https://api.openai.com/v1
- OPENAI_API_KEY=sk-proj-… # billed per token, untraced
+ OPENAI_BASE_URL=https://llm.yourco.internal
+ OPENAI_API_KEY=sk-… # your gateway, your budget, your records
01 The problem
AI usage is growing faster than companies can control it.
Multiple models, multiple vendors, every team holding its own key. No budget ceiling, no per-team breakdown, no record of where prompts go. The invoice arrives as one growing number.
Nobody can say what an agent run costs, which team spent the most, or what the bill looks like next quarter.
No central gateway, no budget ceiling, no rate limit you set. Every laptop and CI runner holds a vendor key.
Switching models means a change programme across every service. So you don't switch — and the vendor knows it.
Legal asks where the prompts go. Nobody has an answer that survives an audit.
02 What we build
One change for your engineers. Full control for you.
Open-source components — LiteLLM, Langfuse, vLLM — wired correctly on infrastructure you own. Typical technologies; final selection depends on workload.
A · ZERO DISRUPTION
Same API, different economics
The gateway answers in the same API shape your tools already speak. Migration is a base URL and a key, not a rewrite.
B · MODEL FREEDOM
Switch models in a routing rule
New release benchmarks well on your evals — it goes live that afternoon. Cheap models take bulk traffic, expensive ones take what needs them.
C · COST VISIBILITY
Cost per team, per call
Every request traced: who asked, which model, how long, how much. Cost per team becomes a dashboard line you hand to finance.
03 What you get back
Cost visibility, per call.
A span waterfall — model, tokens, tool hops, cost — kept as long as your retention policy says. Powered by Langfuse, self-hosted.
How the stack works04 How we work
Assess. Build. Operate. Govern.
The build is the easy half. What matters is the eighteen months after.
05 Why us
Cloud ops since 2015. AI infrastructure since 2023.
Migrations, Kubernetes, Terraform, 24/7 retainers across APAC and China. The discipline — measure, control, operate, report — is the same. The workload is new.
- We run this arrangement in our own production — not a demo built for the pitch.
- China cross-border since 2017 — if your models need to answer inside the mainland, that is an old problem here.
- Two partners work on your engagement directly. You always know who is on the incident.
07 AI infrastructure across jurisdictions
EU, APAC and mainland China — one control plane.
Few European AI infrastructure firms can operate across jurisdictional boundaries. We have been doing exactly that since 2017 — enterprise VPN, site acceleration, ICP filing, cross-border connectivity — and now the payload is AI.
EU workloads route to EU models and infrastructure. China workloads route to approved local providers. Sensitive workloads stay on private inference. The application sees one API; policy decides where the prompt lands.
08 Writing
Nine years of notes from inside other people's infrastructure.
Testing Alibaba Cloud Model Studio (Bailian 百炼)
I’ve written about AI and LLMs a number of times in this blog in the last couple of years, but in all theoccasions I was referring to OpenAI and th...
การทดสอบ Alibaba Cloud Model Studio (Bailian 百炼)
ในช่วงสองสามปีที่ผ่านมาผมได้เขียนเกี่ยวกับ AI และ LLMs หลายครั้งในบล็อกนี้แต่ทุกครั้งที่พูดถึงจะเป็นการอ้างถึง OpenAI และโมเดล GPT3 ของพวกเขา
Xataka Nos Ha Dedicado Un Artículo!
Esto debe ser el cielo, o parecido. Xataka (la publicación sobre tecnología en Español más importante) ha dedicado unartículo a Alberto, creador de...
Migración Entre Regiones De Alibaba Cloud
El cloud, o “la nube”, es algo que tiene un potencial enorme para mejorar la eficiencia de las empresas, pero a menudohace falta hacer ajustes por ...