01 Why this exists
Self-hosted stacks do not decay quietly. They decay expensively.
Six months after a build, the same three things have happened everywhere we look. A better model shipped and nobody moved to it. The routing rules still reflect traffic patterns from launch week. And GPU capacity is sized for a peak that happened once, in March.
None of that is a failure of the architecture. It is what happens when running infrastructure is nobody's actual job — which, in a company whose product is not infrastructure, it correctly is not. The cost shows up as a quiet rise in the invoice, not as an alert.
WHAT WE FIND ON A TAKEOVER
Drift, month six
- Primary model two generations behind
- Routing unchanged since deploy
- GPUs sized for a peak that happened once
- Fallback budget left uncapped after an incident
- Trace retention still on the software default
- Cost per team untracked
Composite of what an unoperated stack looks like half a year in.
02 What's included, every month
03 Tiers
TIER · BASELINE
€1,500 /month
Single-model stack, one or two teams, business-hours coverage.
- Monthly model review
- Quarterly capacity review
- Next business day response
- Monthly written report
TIER · STANDARD — MOST CLIENTS
€3,000 /month
Multi-model routing, org-wide usage, extended coverage across EU and APAC hours.
- Everything in Baseline
- Continuous routing and cost optimisation
- 4-hour response, extended hours
- Per-team chargeback reporting
- Eval pipeline maintained
TIER · CRITICAL
€5,000+ /month
Customer-facing inference, regulated environments, or a mainland China footprint.
- Everything in Standard
- 1-hour response, 24/7 on-call
- Governance and audit support
- Cross-border and residency routing
- Quarterly review with your leadership
Monthly, cancel with 30 days' notice. No minimum term — a retainer you cannot leave is a retainer that stops trying. If the stack is stable and you want to take it in-house, we would rather help you do that than bill you for silence.
04 Taking over someone else's stack
You do not have to have built it with us.
If you already run a self-hosted inference stack that grew organically, we will take it on. First month is an audit: what is actually deployed, what it costs, where it will break, what the quick wins are. You get that written up whether or not you continue.