From Transformer Roles to AI-Enabled Teams

AI-Enabled Teams & Agent Ecosystems — 1 of 13

Whether your team is longing for agents or already working with them, nobody has yet published what a well-run agent team looks like at a firm like yours. This briefing opens a series that treats that gap as an assignment: what changes in each role, what an agent actually consists of, and the first inventory a leader should take.

The question in most leadership meetings is which roles agents will replace.

It is the wrong first move, and not for the comforting reason. It is wrong because it has no answer that anyone can act on. Ask it and you get a debate about capability that ends in either dismissal or alarm, and no decision either way.

The useful question is narrower and harder:

Which work may be supported, delegated or executed within a defined boundary — and who remains responsible when the case is unusual, consequential or wrong?

That question can be answered for one role, this quarter, by people who already understand the work. This series answers it six times, once per Transformer role, and then builds the operating model that holds the answers together.

Where the evidence actually stands

Before this series prescribes anything, you should know what it stands on. We went looking for European firms outside the software industry that run multi-agent setups in production and report honestly on how it has gone. As of September 2026, we did not find one — what exists is single-agent pilots and prototypes at SMEs, first-person accounts from large tech firms, one randomised controlled trial, and a handful of legal rulings. That absence is a finding. It means the operating model for agent teams cannot yet be reported; it has to be authored. That is what this series does, openly, on our evidence standard: published evidence labelled by maturity, our own agent practice with the artefacts to show, professional observation marked “in my work”, and open prescription where the field is still empty.

In my work, I have yet to see anything like a full agent team operating outside software engineering. The closest real arrangements I have observed are engineers collaborating with agents inside a bounded workflow towards a specific outcome — and that makes sense, because software work is project-shaped, with an end point like a deployment. In operational departments the work has no end point: it is ongoing and repeating, and the people doing it rarely think in terms of what could be delegated to an agent. But I have stopped reading that as a mindset gap, because there is a simpler explanation. The tools for building and orchestrating agents are usually handed only to engineers — everyone else never gets the chance to explore what is possible with their own work, while the engineers arrive either for their own productivity or once somebody has already decided a workflow will be overhauled. The untried alternative is the reverse: let the people who live with the pain explore first, as the designers of their own work, and bring engineering in to harden what they find — which is the order the exercise at the end of this briefing assumes. That, more than any gap in the technology, is why the published evidence looks the way it does — and why this series is written for the departments the evidence has not reached yet.

Why “replace the person” is the wrong frame

A role is not a list of tasks. It is a bundle of tasks plus several things that never appear in a job description: a mandate, a set of permissions granted to a named individual, working relationships, standing to accept risk on the organisation’s behalf, and the authority to stop the work when something is wrong.

An agent can take on task bundles. It can take a workflow lane. It cannot hold a mandate. It cannot accept risk on the organisation’s behalf. And when something goes wrong, it cannot be the one a regulator, a customer or a colleague holds responsible. Those are not gaps that a better model closes; they are properties of being an accountable person inside an organisation.

This is no longer only an organisational observation; a tribunal has tested it. When a customer relied on wrong advice from Air Canada’s website chatbot, the airline argued — in writing — that the chatbot was “a separate legal entity that is responsible for its own actions.” The tribunal called the submission remarkable and held the company liable (Moffatt v. Air Canada, February 2024). An organisation tried, before a court, to make an agent the answerable party. The answer was no.

Which means the work in front of a leader is not selection. It is redesign: deciding which parts of a role move, what boundary they move inside, and what the human keeps. Nobody can outsource that to a vendor, because it is a decision about your own accountability.

The six roles do not change in the same way

The Compass groups the people who shape and carry a transformation as Transformers.

A note on the name before anything else. Transformer-based models changed what AI systems can understand, generate, summarize, classify, and reason across. Agents extend that capability into tools, workflows, memory, and action. In the Compass, a Transformer is a person — and the difference between the two is exactly the accountability gap this series is about.

Agents press on each of the six differently, and the difference is where a leader should look first.

  • The Owner finds evidence suddenly cheap and judgement no cheaper. The failure mode is a well-formatted recommendation that nobody actually owns.
  • The Organizer has the widest near-term scope for bounded delegation — routing, status-gathering, chasing, preparing. The failure mode is coordination becoming invisible, so nobody notices it has stopped working.
  • The Creator gets an explosion in the number of options. The failure mode is mistaking more variants for better strategy — and rights, attribution and brand judgement quietly going unowned.
  • The Implementer can change systems faster. The failure mode is an action that reaches production without a change authority, and cannot easily be undone.
  • The Tester can generate tests trivially. The failure mode is losing independence: a system grading its own work is not assurance.
  • The Maintainer inherits a genuinely new burden — running, watching and stopping something that acts on its own. The failure mode is that nobody was given this job at all.

Read across those and a pattern appears. The task shifts differ; the thing at risk is the same in all six. It is always the point where someone was supposed to be answerable.

What an agent actually is

Leaders are handed the word “agent” as though it named one thing. It names an assembly of nine, and a leader who can name them can ask useful questions without becoming an engineer.

  1. Model — the reasoning engine. Interchangeable, and the least interesting part of the design.
  2. Instructions — the standing brief: what this agent is for, how it should behave, what it must never do.
  3. Tools — what it can reach and act on: a search, a database, a ticketing system, an email account.
  4. Knowledge and context — the material it is allowed to read, and which of it counts as authoritative.
  5. Memory and state — what it carries between steps and between sessions, which is also what it can carry between people who should not share information.
  6. Permissions — the identity it acts under. An agent inherits whatever that identity can do, including things nobody intended.
  7. Evaluation — how anyone knows it is still good, on a defined set of cases, after something changes.
  8. Monitoring — how anyone sees what it actually did, after the fact.
  9. Human hand-offs — where it stops, who it stops to, and what that person receives.

Almost all leadership attention goes to the first item. Almost all the risk sits in the other eight. When a supplier presents a capability, these nine are the questions that turn the presentation into something you can govern: not “how good is it?” but “what can it reach, under whose identity, judged how, watched by whom, and stopping where?”

And on the seventh component — evaluation — one number is worth carrying into every meeting where somebody reports how well an agent is working. In the only randomised controlled trial of its kind to date, experienced developers using AI tools were measured 19% slower on real tasks in their own repositories, while believing afterwards that they had been about 20% faster (METR, July 2025). METR itself treats the result as a snapshot of early-2025 tools; the durable finding is the thirty-nine-point gap between what people felt and what the clock said. Evaluation means measurement against defined cases. Asking users whether the agent helps is not evaluation; it is a satisfaction survey.

Three patterns: one, a suite, or a lane

Nearly every proposal you will see is one of three shapes. Telling them apart is the leader’s job, not the architect’s: the shape decides what you will be asked to build, fund and answer for. Everything else in this series starts from that call.

One constrained assistant. A single agent with a clear brief and a few tools, used by a person who remains in the work: drafting, summarising, comparing, preparing. The human does the job; the agent shortens parts of it. This is augmentation, and it is where most real value currently sits.

A role-specific suite. Several specialised agents with separate responsibilities around one role — for a Tester, perhaps a test generator, a regression runner, a data-quality checker and an evidence collector. Each has its own brief and its own boundary. The human coordinates and decides. Value comes from separation of purpose; cost comes from having several things to maintain, evaluate and watch.

A bounded workflow lane. A repeatable process where steps hand off to each other, some executed without a person in the middle. This is the only one of the three that is meaningfully “agentic”, and the only one that can act while nobody is looking. It demands the most: explicit routing, stop rules, escalation, logging and a named operating owner.

The engineering literature is consistent on the point leaders most need to hear: prefer the simplest pattern that meets the need. Anthropic’s own guidance for builders makes the same argument — more agents introduce cost, latency, coordination failure and diffuse accountability, and a single well-scoped call or a deterministic workflow is frequently the better answer (Building Effective AI Agents). Complexity is something to justify, not something to aim at.

The redistribution is measurable where agents already run at volume. Dropbox’s internal agent platform now produces roughly one in twelve of the company’s pull requests — and its engineering leadership’s own conclusion, published May 2026, is not a story about replaced engineers. Acceleration moved the constraint downstream: review queues, testing, validation, release coordination and operations came under strain, while engineers stayed accountable for intent, architecture, quality and release (Beyond code generation — Dropbox Tech, May 2026). More agent execution did not remove the human work. It changed which work the humans are the bottleneck for.

Where a proposal sits on that scale also decides how much of this series applies to it. If what your team has is one assistant, most of what follows is light. If someone is proposing a lane that acts on its own, all of it applies. Navigating the Black Box Continuum sets out the underlying distinction between automation, assistance and genuine agency.

The leader’s first four maps

Before any of this can be decided, four things have to be visible. They are not documents for their own sake; they are the four things you would need in order to explain the arrangement to someone who was not in the room.

  1. The work map. What actually happens in this role, step by step, including the hand-offs, waits and exceptions. Not the process document — the work, described as situations a person outside the team would recognise. Work scoped as a category cannot be evaluated, and what cannot be evaluated cannot be stopped. Map the Work Before You Automate It is the method.
  2. The decision-rights map. Who may decide what, who may approve, who may override, and who may stop the work entirely.
  3. The data and evidence map. What may be read, what counts as authoritative when sources disagree, and what must never enter the system at all.
  4. The escalation map. What must reach a person, which person, and what they receive when it does.

The third map is the one small firms get right when they take it seriously. InVIDO, a German custom-furniture manufacturer, built a knowledge-assistant prototype with the data and evidence map as the design itself: the firm supplied the documents, answers had to be bound to those sources rather than to the model’s general knowledge, people who knew the material assessed answer quality, and self-hosting was a hard requirement because sensitive company data was not going to an external cloud provider (Mittelstand-Digital Zentrum Chemnitz, 2026). It is a prototype, built to decide whether the long-term investment is justified — and that is the point: none of those four decisions needed an engineer. All of them needed an owner to insist.

Most organisations have none of these written down for a role that is already using agents. That is not a governance failing so much as a sequencing one: the tools arrived faster than the maps, usually from the bottom up, and often without anyone deciding. The fix is not to stop that exploring — people testing agents against their own pain is the best signal you will get about where agents belong. The fix is to give it the maps, so that quiet drift becomes designed exploration.

Start here: classify one role’s work

The practical first move is small enough to do in a workshop. Take one role. List its recurring work — twenty items is plenty. Then sort each item into one of four levels:

  • Human-led — a person does this, possibly with better information than before.
  • Augmented — an agent assists; the person still does the work and owns the output.
  • Delegated with approval — an agent prepares or proposes; a named person approves before anything takes effect.
  • Bounded autonomous — an agent acts inside an explicitly defined lane, with stop rules and a monitored boundary.

Then, for each item you did not mark human-led, answer two questions. What happens if this is wrong? And can we undo it? Anything irreversible or consequential moves up a level towards the human, regardless of how capable the system is. This is the point where authority and capability get separated, and separating them is the whole discipline. Delegate, Automate, or Keep Human: The Agent Handoff Map is the working technique for doing this step by step.

The classification is not a consultant’s abstraction; small firms are building it directly into systems. EMES Kabelbaum Konfektion, a Saxon switchboard manufacturer with around 150 regular customers, designed an email-to-ERP order-entry pipeline around exactly these levels, expressed as a traffic light: green orders book themselves into the ERP, yellow orders stop for a person’s review, red orders are entered manually, and corrections feed back into the system (Mittelstand-Digital Zentrum Chemnitz, 2026). The pipeline is only now entering a proof of concept, and two decisions made before any of it runs are worth copying. The levels were set by the people who know the work, not by a vendor. And the firm’s first AI decision was a refusal — a cloud-based calculation tool was rejected because margins and hourly rates do not leave the company. The boundary came before the capability.

What you are left with is an inventory: this role’s work, the level each part sits at, and who carries the consequence. It is the artefact the rest of this series builds on, and it is worth doing badly and quickly rather than perfectly and never.

The inventory is also what makes agents work at all. FundFrame, a small Copenhagen investment-software team, attributes its agent-driven delivery speed less to the models than to operating work that looks exactly like this: work shaped into well-scoped items before any agent touches it, durable written context agents can navigate, and — once generation got fast — testing emerging as the new bottleneck, answered with automated end-to-end checks and independent review, including deliberately using different models to check one another’s output (Our Agentic Development Setup — FundFrame, March 2026). Shaping, context and review are not prompt tricks. They are the operating work, and they are led, not delegated.

In my work, I have yet to see a team run this exercise before the tools arrived. So we ran it on ourselves — and honesty requires stating the scale. PathPatron is produced with several agents: some support research, drafting and editing; others update the website. They are not connected to each other. I am the interconnection: I review what one produces, then hand the result — usually a plain markdown file — to the next. One agent runs on a Mac mini; browser-based tools carry the rest. Every recurring task sits at a declared level: factual corrections run inside a pre-approved changeset with a verification checklist; anything touching voice or claims is gated on my sign-off, article by article; anything first-person stays human-led, full stop. By this briefing’s own classification, that is a suite with a person at every seam — mostly delegated with approval, nothing autonomous, nothing acting unseen. The evidence for firms like yours does not exist yet. But the method does, it runs on hardware you already own, and it starts with the people who know the work.

What this series does

Thirteen briefings, in three movements: two that orient, six that take the Transformer roles one at a time — Owner, Organizer, Creator, Implementer, Tester, Maintainer — and five that build the operating model around them. The series ends with the Agent Team Blueprint v1.0: the whole arrangement on paper, versioned and dated, with the parts most likely to be proven wrong marked as such — because where evidence does not yet exist, an explicit standard is the honest alternative to pretending it does. The full route, reading order and publication status live on the Path page: AI-Enabled Teams & Agent Ecosystems. If you run any of it in your organisation, tell us what broke; anonymised reports from practice are how the missing evidence gets built.

That is the PathPatron stance. Agents do not replace roles. They redistribute work across a boundary that somebody has to draw — and drawing it, then holding it, is the leadership job that does not delegate.