InsightsAI Strategy & Transformation
← All insights All tools

Insight

Operating Model: Stop Running Agents Like Side Projects

Published 2026-05-10 · Mike Kennedy

Agentic EnterpriseStrategyReal People ImpactChange Management & CultureOperating Model & GovernanceLeadership

Part 3 of 9 — Operating Model

A CIO I respect once told me, "We have 200 people working on AI." I asked who they reported to. The answer involved seven boxes on an org chart and three matrix relationships. That is not an operating model. That is a hostage situation.

Most enterprises have an AI strategy long before they have an operating model. The strategy gets approved in a boardroom. The operating model gets cobbled together in the hallway after, by whoever happens to be there. Predictably, the cobbling does not scale.

This article is about the three capabilities that make up the Operating Model focus area in the agentic maturity framework, and the unglamorous work Aurelia did to stop running agents like side projects across a portfolio that included theme parks, a streaming service, a cruise line, an airline, and a film studio.

The Three Capabilities of an Agentic Operating Model

Operating Model has fewer capabilities than Strategy, but they are heavy ones:

  1. Organizational Design — the AI-related job profiles, roles, accountability structures, and reporting lines across teams

  2. Pilot Strategy — how pilots are scoped, run, measured, and retired or scaled

  3. Agentic Supervision — the processes that explain, trace, and verify agent decisions, plus the discipline to monitor for drift

Each of these is a muscle most enterprises do not have. Building them takes 12 to 18 months. Pretending you have them when you do not is one of the more expensive forms of self-deception in this field.

What Level 1 Looks Like

Low-maturity operating model has a particular signature:

If three of those describe your enterprise, you are running agents like side projects, and the side project is going to bite you.

Aurelia's Starting Point

When I arrived at Aurelia, the AI organization was 187 people scattered across 11 business units, reporting through 23 different managers, with no shared operating cadence. The pilot success rate — defined as "made it from pilot to production with measurable impact" — was 6%. Six percent.

That number, when the CDO finally surfaced it to the executive committee, was the moment the operating model conversation became unavoidable. The Studios president had been quietly assuming his pilots were succeeding. The Parks president had been assuming the same. The aggregated reality was harder to ignore.

The Big Rocks Aurelia Had to Move

Rock 1: Funding a Center of Excellence without killing distributed innovation

This is the eternal tension. Centralize too hard and you starve the business units of agency. Distribute too hard and you get the 187-people-23-managers problem. Aurelia's answer was a hub-and-spoke model: a deep COE at the center (about 90 people) plus embedded pods inside each major business unit (4-7 people each, with the largest pods inside Parks & Resorts and Aurelia+).

The COE owned: the platform, the standards, the eval framework, the supervision discipline, and the org-wide capability roadmap. The embedded pods owned: the use case, the BU relationship, the change management, and the Real People Impact at the front line.

The interface between them was governed by a set of operating principles that fit on one page. The most important principle, in my experience: the COE is accountable for capability; the pod is accountable for outcome. If a pod fails because the platform let them down, that is a COE failure. If the platform works and the pod still misses, that is a pod failure. This boundary saved more relationships than any RACI chart in the building.

A meaningful design choice: the embedded pods at Aurelia were physically located inside the BUs they served. The pod inside Aurelia Skies sat at the airline's operations center. The pod inside Aurelia Voyages was based at the cruise line's home port. The pod inside Aurelia Pictures sat on the studio lot. Proximity matters. The pods that tried to support a BU from a central campus consistently underperformed the pods that lived inside the business they served.

Rock 2: Hiring AI Product Managers

This is a role most enterprises do not have, and most enterprises do not realize they do not have it. An AI Product Manager is not a traditional PM with an "AI" label. They have a different skill stack: they understand model behavior at a working level, they can read an eval report, they can write a prompt that holds up, they can negotiate trade-offs between latency and accuracy, and they can translate between data scientists and BU leaders without losing both audiences.

Aurelia needed 28 of them. They hired 9 externally and grew 19 internally through a 9-month rotation program. The rotation program turned out to be the more durable approach. Internally grown AI PMs already understood the business — they knew the difference between a Studios production schedule and a Live tour schedule, they knew why a cruise sailing's daily program was not the same as a hotel's daily program, they knew why a guest's expectations at the Kingdom flagship were different from a guest's expectations at Aurelia Riviera. The hard part — model intuition — could be taught. The other direction was harder.

Rock 3: Building the supervision discipline

This was the rock most companies skip, and Aurelia nearly skipped it too. Agentic Supervision is the operating muscle that lets you understand why an agent did what it did, verify whether it was right, and detect when it has started doing something different than it used to do.

Aurelia stood up an 8-FTE Agentic Supervision team inside the COE. Their job was, plainly, to watch the agents. They built traceability tooling that captured every agent decision with its inputs, retrieved context, model output, and downstream action. They built a weekly drift report that compared agent behavior in the current week against a baseline. They built an escalation path for when an agent's behavior shifted in a way the team couldn't immediately explain.

Eight FTEs across 11 BUs and what would eventually become 60+ production agents. That is what stood between Aurelia and a class of incident that would have made the trades and probably the front page.

What "Leading" Looks Like

Eighteen months in, Aurelia's operating model looks like this:

This is not glamorous. It is, however, what stops the program from collapsing under its own weight when scale arrives — and in a 24/7/365 operation that never closes its parks, never grounds its fleet, and never goes dark on its streams, scale arrives fast.

The RPI at the Other End

The pilot success rate matters for one reason: every failed pilot is a team of humans who put six months of work into something that died. They learn cynicism. They become harder to recruit for the next pilot. They tell their friends. They leave.

When Aurelia moved the pilot success rate from 6% to 71%, the second-order effect was on the people doing the work. AI engineering attrition dropped by half. The internal Net Promoter Score for "would you work on another agent project at Aurelia" went from -22 to +41. People started self-selecting into this work instead of being assigned to it. The embedded pod inside Aurelia Live, which had been considered a hardship assignment 18 months earlier, became one of the most requested rotations in the program — partly because the work was meaningful and partly because the pod was actually shipping.

That is the operating model RPI. Not just outputs. Whether the humans inside the program are running toward it or away from it.

What to Do This Week

If you want to test where your operating model sits, try this:

  1. Pull the list of every agent or AI initiative currently active in the company.

  2. For each one, identify the named individual whose career performance is tied to the outcome of that initiative. Not a steering committee. Not a department. One person.

  3. Count how many of your initiatives have a real answer to that question.

If the answer is fewer than half, your operating model is not yet the load-bearing structure you need it to be.

Next article: Expertise & Culture. The bottleneck no one wants to talk about, and the day Aurelia's most beloved Principal Imagineer walked out the door.


This is Part 3 of a 9-part series on agentic enterprise maturity.