Services · Managed Run & Optimisation

Launch is the easy part.

We run the systems we (or others) have built. Under an SLA, with the metrics in plain sight.

A live AI agent isn’t a finished project.

Models change, prompts drift, costs creep, integrations move, regulations evolve and operations expose edge cases the test set never covered. Managed Run is the practice of keeping these systems honest, fast, accurate and affordable. Every day, not just on launch day.

The systems we build

What it covers

The unglamorous work that keeps it trustworthy.

01

Operational monitoring

Health checks, uptime, latency, error rates, queue depths. The boring, essential plumbing: measured, alerted, and on-call when it matters.

02

AI-specific evaluations

The eval harness we built at launch - or that we’ll build if you’ve inherited a system without one - runs on every change and on a schedule in production. We watch for quality regressions in plain language as carefully as we watch for HTTP 500s.

03

Model and prompt management

Frontier models change every few months; open-weights models change every few weeks. We track what’s available, test it against your evals, and recommend - and apply - upgrades that move quality or cost in the right direction. Prompts and configurations are versioned, reviewed and rolled out with the same care as code.

04

Cost optimisation

Token spend, infrastructure cost, vendor fees: visible in one place, trended, and managed. We cut AI run-costs by routing, caching, batching and right-sizing models. We share the dashboard.

05

Incident response

When something breaks, you have one number to call and one team accountable. Post-incident, you get a real review: what happened, why, what we changed to make it less likely next time. We don’t hide behind vendors.

06

Continuous improvement backlog

A live backlog of small, high-value improvements, based on the metrics, on user feedback, on the evals, on what we and your operations team are seeing. We run it like an engineering backlog, not like a steady-state support queue.

What we run

Anything we’ve built, and quite a lot we haven’t.

  • AI agents and agentic workflows
  • Automated business processes: Camunda, Temporal, Power Automate, Nintex, custom
  • Custom AI applications: web, mobile, internal
  • Legacy workflow estates being progressively modernised
  • Integrations and data pipelines feeding any of the above

If it’s part of how AI does work in your organisation, we can take it on.


How it works

Onboard, operate, improve.

01

Onboarding

We inventory the systems, the integrations, the data flows, the runbooks (if any), the eval suites (if any), and the people who care about each service. We close obvious gaps before we go live: telemetry that’s not wired up, missing runbooks, evaluations that don’t exist, secrets that are everyone-can-read.

2-4 wks
02

Steady state

We work to an agreed SLA, with a single accountable owner on our side. Weekly metrics and backlog review with you, monthly business review with sponsors. Quarterly architecture review.

Ongoing
03

Continuous evolution

This isn’t a hands-off support contract. We expect to be improving the service every month, and we publish what we did.

Monthly

Service levels

Tailored to the criticality of each service.

Tier Target uptime Critical-incident response Quality eval frequency
Standard 99.5% 4 business hours Weekly
Business critical 99.9% 1 hour, business hours Daily
Mission critical 99.95% 30 minutes, 24/7 Continuous

Typical tiers. SLAs are agreed per service during onboarding.

In plain words

What you get from us.

Entry point
2-4 week onboarding, then monthly
Output
Operating service, monthly report, improvement backlog
Team
Service lead + AI engineer + ops engineer + escalation to practice leads
Term
Rolling monthly after initial 6-month term
All services
  • A single person you can call when something feels off.
  • A monthly statement of what we did, what we changed, what it cost, and what’s queued up next.
  • An honest answer when something we changed didn’t work.
  • A roadmap of improvements that’s actually being delivered, not just documented.
  • Confidence that the system you bet on at launch is the system running three years later.

Got something in production you’d rather not run alone?

We’ll take a look. Tell us what you’ve shipped, where it’s running and what’s keeping you up at night.