Skip to content

Guztia AI Infra Ops

Guztia /Services /AI Ops

SVC-02 · Monthly retainer

AI Ops Retainer

The build is the easy half. This is the other half — model rotations, routing and cost tuning, GPU capacity, incidents at unsociable hours, and a monthly report that makes your AI spend legible to someone who does not read YAML.

01 Why this exists

Self-hosted stacks do not decay quietly. They decay expensively.

Six months after a build, the same three things have happened everywhere we look. A better model shipped and nobody moved to it. The routing rules still reflect traffic patterns from launch week. And GPU capacity is sized for a peak that happened once, in March.

None of that is a failure of the architecture. It is what happens when running infrastructure is nobody's actual job — which, in a company whose product is not infrastructure, it correctly is not. The cost shows up as a quiet rise in the invoice, not as an alert.

WHAT WE FIND ON A TAKEOVER

Drift, month six

  • Primary model two generations behind
  • Routing unchanged since deploy
  • GPUs sized for a peak that happened once
  • Fallback budget left uncapped after an incident
  • Trace retention still on the software default
  • Cost per team untracked

Composite of what an unoperated stack looks like half a year in.

02 What's included, every month

Model rotation
New open-weight releases benchmarked against your evals and your traffic, not a leaderboard. If a model wins, it rolls out behind a canary. If it loses, you get a paragraph saying so. This happens roughly monthly at the current release cadence.
Routing & cost tuning
Continuous rebalancing across cost, quality and latency. Cheap models absorb what they can handle, expensive paths stay for what needs them, caching catches the repeats. Most months this is where the retainer pays for itself.
GPU capacity
Right-sizing against measured utilisation, reserved versus on-demand versus spot decisions, autoscaling that actually scales down, and quota requests raised before you need them rather than during the incident.
Observability ops
New teams and projects onboarded, dashboards built for questions people are actually asking, eval pipelines maintained, retention enforced, storage kept from growing without bound.
Incident response
OOMing nodes, driver regressions after an update, latency cliffs under load, fallback paths that silently became the primary path. Response targets below; escalation goes to a named engineer, not a queue.
Upgrades
Serving engine, gateway, observability, CUDA, Kubernetes. Tested in staging, rolled out in a window you pick, rolled back when they misbehave — which, with inference stacks moving at this pace, some of them will.
Monthly report
Cost per team, project and model. Token volume and its trend. Latency percentiles. Eval scores over time. What changed, what broke, what we recommend next. Written so your CFO can read it without a translator.
Access to a human
A shared channel with the engineer who built your stack. Not a ticket portal, not a rotating account manager.

03 Tiers

TIER · BASELINE

€1,500 /month

Single-model stack, one or two teams, business-hours coverage.

  • Monthly model review
  • Quarterly capacity review
  • Next business day response
  • Monthly written report

TIER · STANDARD — MOST CLIENTS

€3,000 /month

Multi-model routing, org-wide usage, extended coverage across EU and APAC hours.

  • Everything in Baseline
  • Continuous routing and cost optimisation
  • 4-hour response, extended hours
  • Per-team chargeback reporting
  • Eval pipeline maintained

TIER · CRITICAL

€5,000+ /month

Customer-facing inference, regulated environments, or a mainland China footprint.

  • Everything in Standard
  • 1-hour response, 24/7 on-call
  • Governance and audit support
  • Cross-border and residency routing
  • Quarterly review with your leadership

Monthly, cancel with 30 days' notice. No minimum term — a retainer you cannot leave is a retainer that stops trying. If the stack is stable and you want to take it in-house, we would rather help you do that than bill you for silence.

04 Taking over someone else's stack

You do not have to have built it with us.

If you already run a self-hosted inference stack that grew organically, we will take it on. First month is an audit: what is actually deployed, what it costs, where it will break, what the quick wins are. You get that written up whether or not you continue.

Book a takeover audit

Month 1
Audit and stabilise
Month 2
Quick wins — usually cost
Month 3+
Steady operation

Own the stack before the bill owns you.

Thirty minutes. Bring your current AI invoice. We look at what your teams send to model APIs, what it costs today, and what a self-hosted stack would cost instead. No deck, no discovery phase.