All case studies

Confidential engagement

Sector IT services

Engagement AI Build · AI Consultancy

Duration Couple of months

Published with the client's permission. Details of the business have been generalised at their request; the system, the build, and the results are unchanged.

A multi-agent system: 30+ agents coordinating the work their teams did by hand

Agents running

30+

Hours back / week

10+

Running since

April 2026

1

The business

An IT services firm whose internal delivery workflow had outgrown what people could carry by hand. The work ran across several internal teams, each carrying a piece of the same process and none of them seeing the whole.

2

The problem

A recurring workflow that touched multiple teams, each doing the same narrow job by hand, pulling information from their own sources, reformatting it, passing it along. The work scaled with the business, and the coordination cost grew faster than the work itself. What made this different from a routine build was the shape of the answer: not one pipeline, but a multi-agent system with 30+ agents coordinating across the teams.

3

What we found

Scoping revealed that the work split naturally into agent teams rather than a single flow. Each team owned one responsibility, but the boundaries between them, and the shared state they all depended on, were where the real complexity lived. That's why the engagement ran as AI Build with AI Consultancy alongside it: the architecture decisions mattered as much as the code.

4

What we built

  • Agent teams, each owning one responsibility: intake, synthesis, quality checks, routing, drafting. Thirty-plus agents in total, working in parallel rather than one long chain, so a failure in one stream doesn't stop the others.
  • A shared orchestration layer that routes work between teams, holds the state they all read and write, and decides when a task needs a person in the loop.
  • Human-in-the-loop gates at the points where judgement genuinely matters. The system drafts, a person approves, only then does it go out. Most of the work flows through unattended; the exceptions land exactly where they should.
  • An evaluation harness in place from the first week, so every change to prompts or agent behaviour is measured against the same baseline rather than judged by feel.
  • What we deliberately left alone: the handful of steps that needed a named human's judgement and weren't going to get faster with automation.
5

What changed

One system where there used to be a hand-off between teams. The work now flows through the agent teams and surfaces to a person only where judgement genuinely doesn't transfer. Hours recovered, turnaround, and error rate are all reported to the same baseline the system evaluates against.

6

What was harder than expected

The hardest part was the shared state. Each team had its own idea of what a unit of work looked like, and the orchestration layer had to reconcile them before the agents could trust it. Invisible in a single-agent demo, decisive in a multi-agent system.

7

Client quote

We'd been carrying this workflow for years and treating it as the cost of doing business. Now a system of 30+ agents runs the flow, and people only step in where judgement matters. We can measure the hours back every week. The write-up was specific enough that our own team could have built it.

· Head of Operations, IT services firm

Get started

Recognise any of this in your own business?

See how the audit works