Skip to content

Guztia AI Infra Ops

Guztia /Services /Why Self-Host

Cost · Control · Forecast

Why Self-Host AI

Per-token API pricing feels convenient until the invoice becomes the strategy. Here is an honest comparison of vendor APIs versus self-hosted models — what you save, what you give up, and when the hybrid answer is the right one.

01 The API pricing model

You are renting inference at someone else's margin.

Per-token pricing is simple to start. Every engineer, CI job and agent gets a key. The bill arrives as one number. Nobody can forecast next quarter because nobody can see which team spent what, or what would happen if list prices rose.

Frontier labs have been selling tokens and seats below what the GPUs cost them, funded by capital rather than revenue. That gap closes — through price rises, tighter limits, or both. The teams that built on a subsidy and never measured it will find out the hard way.

Unpredictable
Volume grows with every new tool and agent. The invoice grows with it, after the fact.
Opaque
One line item. No chargeback. No answer when legal asks where the prompts went.
Sticky
Switching models means a change programme across every laptop and pipeline — so you don't.
Subsidised
Today's price is not tomorrow's cost. Plan for the day the subsidy ends.

02 The self-hosted model

You buy GPU-hours and divide.

Self-hosting flips the unit of cost. Instead of paying per token at the vendor's list price, you pay for capacity — reserved or on-demand GPUs — and utilisation determines what each million tokens actually costs.

Teams spending €10k+/month on API calls typically see 50–70% reduction once bulk traffic moves to open-weight models on their own cloud, with a capped vendor path kept for the hard cases. That is an illustration, not a quote — your traffic shape decides the real number.

The build itself is a fixed fee (€8k–€25k depending on scope). GPU and cloud costs stay on your provider bill. There is no Guztia margin on your tokens.

ILLUSTRATIVE · NOT A QUOTE

What changes on the invoice

Before
€10k–€80k/mo API spend · one invoice · no per-team breakdown
After
GPU + cloud on your account · gateway budgets · cost per team traced
Build
€8k–€25k one-time · typically pays back within a quarter at mid-volume
Ops
€1.5k–€5k/mo retainer if you want someone else on the pager

03 What you give up

We are not going to sell you a free lunch.

Self-hosting is not always the right call. At low volume, the operational overhead and GPU floor cost can exceed what you pay OpenAI or Anthropic. Frontier models still win on long-horizon agents, hard novel reasoning, and some multimodal work. Pretending otherwise costs more than the licence fees you saved.

Ceiling
Open weights have closed most of the gap for enterprise bulk work. Not all of it. Keep a capped vendor route.
Ops
Someone has to rotate models, tune routing, and answer when a node OOMs at 3am. That is either you or a retainer.
Upfront
A fixed project fee and a few weeks of attention. Not free. Often cheaper than another quarter of surprise invoices.

WE WILL SAY NO

Sometimes the assessment ends with "not yet"

If your volume does not justify GPUs, we say so and you keep the analysis. Often the first step is still worth it: put a gateway in front of the vendor APIs so you can finally see what you spend — and move later, without another migration project.

04 The hybrid answer

Route by difficulty, not by loyalty.

A self-hosted mid-size model takes the volume. A larger self-hosted model takes what it cannot handle. A frontier API, behind a hard budget cap, takes what neither can — and every call is traced, so you see exactly what the escape hatch costs and whether it is shrinking.

In most deployments we run, it shrinks. Engineers change two environment variables. Finance gets a forecast. Legal gets an audit trail. That is the whole point.

Bulk
Self-hosted open-weight · predictable GPU cost
Hard cases
Larger self-hosted or capped frontier API
Visibility
Every call traced · cost per team · budgets that stop spending
Exit
Open source · your account · no Guztia component to remove

Own the stack before the bill owns you.

Thirty minutes. Bring your current AI invoice. We look at what your teams send to model APIs, what it costs today, and what a self-hosted stack would cost instead. No deck, no discovery phase.