01 The API pricing model
You are renting inference at someone else's margin.
Per-token pricing is simple to start. Every engineer, CI job and agent gets a key. The bill arrives as one number. Nobody can forecast next quarter because nobody can see which team spent what, or what would happen if list prices rose.
Frontier labs have been selling tokens and seats below what the GPUs cost them, funded by capital rather than revenue. That gap closes — through price rises, tighter limits, or both. The teams that built on a subsidy and never measured it will find out the hard way.
02 The self-hosted model
You buy GPU-hours and divide.
Self-hosting flips the unit of cost. Instead of paying per token at the vendor's list price, you pay for capacity — reserved or on-demand GPUs — and utilisation determines what each million tokens actually costs.
Teams spending €10k+/month on API calls typically see 50–70% reduction once bulk traffic moves to open-weight models on their own cloud, with a capped vendor path kept for the hard cases. That is an illustration, not a quote — your traffic shape decides the real number.
The build itself is a fixed fee (€8k–€25k depending on scope). GPU and cloud costs stay on your provider bill. There is no Guztia margin on your tokens.
ILLUSTRATIVE · NOT A QUOTE
What changes on the invoice
03 What you give up
We are not going to sell you a free lunch.
Self-hosting is not always the right call. At low volume, the operational overhead and GPU floor cost can exceed what you pay OpenAI or Anthropic. Frontier models still win on long-horizon agents, hard novel reasoning, and some multimodal work. Pretending otherwise costs more than the licence fees you saved.
WE WILL SAY NO
Sometimes the assessment ends with "not yet"
If your volume does not justify GPUs, we say so and you keep the analysis. Often the first step is still worth it: put a gateway in front of the vendor APIs so you can finally see what you spend — and move later, without another migration project.
04 The hybrid answer
Route by difficulty, not by loyalty.
A self-hosted mid-size model takes the volume. A larger self-hosted model takes what it cannot handle. A frontier API, behind a hard budget cap, takes what neither can — and every call is traced, so you see exactly what the escape hatch costs and whether it is shrinking.
In most deployments we run, it shrinks. Engineers change two environment variables. Finance gets a forecast. Legal gets an audit trail. That is the whole point.