Skip to content

Scaled agile

When an AI program grows past one team, "how do we deliver?" becomes "how do many teams deliver one AI product?" — and two families of answers dominate: SAFe (Scaled Agile Framework — prescriptive, role-rich, portfolio-aware) and LeSS (Large-Scale Scrum — minimal, de-scaling, feature-team-first). They resolve the scaling problem in opposite directions.

This page presents SAFe as the primary framework and LeSS as the peer, as linked method tabs — pick SAFe once and every tabbed block on the page follows.

Sources verified 2026-08-09

The AI-program lens

Three things scaling an AI program surfaces that a generic "scale-agile" text doesn't linger on:

  • The platform-team vs feature-team tension is unusually acute. ML platform (feature store, training infra, serving, MLOps) is a classic Team Topologies platform team; feature teams that ship models often need its capabilities on their timeline. Whichever framework you pick has to answer how does platform capacity get prioritized against feature-team asks?
  • The single-backlog question is a governance question. Multiple product teams shipping models against a shared Responsible-AI policy needs one backlog — or a mechanism that behaves like one. Otherwise the governance items (fairness audits, drift-monitoring build-outs, guardrails work) get perpetually deprioritized against feature work.
  • The MLOps CI/CD pipeline must include CT — Continuous Training. A "definition of done" for a model iteration that matches software's DoD is a leading indicator that the org will accumulate drift debt.

Both frameworks below are used through this lens.

The two frameworks

SAFe is a prescriptive, role-rich framework with four configurations you compose upward:

  • Essential SAFe — Team + Program level (a single Agile Release Train — ART).
  • Large Solution SAFe — adds coordination across multiple ARTs building one solution.
  • Portfolio SAFe — adds strategic investment, Lean budgets, and portfolio value streams.
  • Full SAFe — all four levels.

The four levels — layered

flowchart TB
  subgraph Portfolio["Portfolio (Portfolio SAFe)"]
    LPM[Lean Portfolio Management<br/>Strategic themes, Lean budgets,<br/>Portfolio Kanban, WSJF prioritization]
    VS[Value streams:<br/>Development value stream<br/>+ Operational value stream]
  end

  subgraph LargeSolution["Large Solution (Large Solution SAFe)"]
    STE[Solution Train:<br/>Solution Train Engineer,<br/>Solution Architect, Solution Mgmt]
  end

  subgraph Program["Program / ART (Essential SAFe)"]
    ART[Agile Release Train:<br/>50 -- 125 people, 8 -- 12 weeks / PI]
    PI[PI Planning event<br/>every PI]
    RTE[RTE + Product Mgmt<br/>+ System Architect]
  end

  subgraph Team["Team (Essential SAFe)"]
    SCRUM[Scrum / Kanban teams:<br/>PO + Scrum Master + team,<br/>2-week iterations]
  end

  Portfolio --> LargeSolution --> Program --> Team

  classDef pnode fill:#e3f2fd,stroke:#1565c0,color:#000
  classDef lnode fill:#fff3e0,stroke:#e65100,color:#000
  classDef prognode fill:#e8f5e9,stroke:#2e7d32,color:#000
  classDef tnode fill:#f3e5f5,stroke:#6a1b9a,color:#000
  class LPM,VS pnode
  class STE lnode
  class ART,PI,RTE prognode
  class SCRUM tnode

AI-adjusted SAFe — what each level actually delivers

Level Standard SAFe outcome AI-program adjustment
Team 2-week Scrum/Kanban iterations delivering user stories Each team's Definition of Done for a model iteration includes eval-set pass + model card update + drift-monitor hook. Not just "story merged."
Program / ART An ART of 50–125 people delivers one program on a shared PI cadence (8–12 weeks) with PI Planning at start The program is a product line (e.g. "recommendations") — feature teams that own end-to-end model+integration, sharing an ML platform ART's outputs. PI objectives are tied to model milestones ("model v2 hits KPI Δ ≥ +3% by end of PI") not just features shipped.
Large Solution Multiple ARTs coordinate to deliver one solution too big for one ART An AI Platform ART (feature store, training infra, serving, MLOps CI/CD/CT, guardrail plane) delivers shared capabilities on the same PI cadence as the product ARTs. Solution Architect covers the ML-platform architecture.
Portfolio Lean Portfolio Management funds value streams via Lean budgets, prioritizes via WSJF (Weighted Shortest Job First) Value stream mapping treats data pipelines and model lifecycle as first-class Development Value Streams. WSJF for AI investments folds in risk of not doing the responsible-AI work on the cost-of-delay side. Enablers (data quality, drift monitoring, MLOps CI/CD/CT capacity) are funded as first-class Portfolio Epics, not squeezed inside feature-team capacity.

PI Planning as the AI-alignment event

PI Planning is the recurring 2-day face-to-face (or virtual) ceremony every 8–12 weeks where all ART teams together commit to PI objectives. For an AI product ART, PI Planning is where the ML platform ART's promised capabilities meet the feature teams' model roadmap. Concrete artefacts a well-run AI PI Planning produces:

  • Feature-team PI objectives that name the model KPI Δ they'll try to hit and the eval methodology.
  • Explicit dependencies drawn from feature teams onto the AI Platform ART (e.g. "feature team X needs feature-store schema Y in place by iteration 3").
  • Enabler work on the roadmap for guardrails, drift monitoring, fairness audits — not just feature work.
  • A CoP (Community of Practice) for ML engineers across teams — the platform-adjacent cross-team knowledge sharing that stops each team reinventing the eval harness.

Continuous delivery, adjusted

SAFe's "Continuous Delivery Pipeline" is the ART's flow from idea → deploy. For AI programmes it must be extended with a CT — Continuous Training — stage:

Continuous Exploration -> Continuous Integration ->
  Continuous Deployment -> Continuous Training (retrain
  on drift) -> Release on Demand

Without the CT step your CI/CD pipeline can promote code, but the model it serves will drift out of currency between promotions. MLOps discipline is what turns "CI/CD" into "CI/CD/CT."

LeSS — Large-Scale Scrum, by Craig Larman and Bas Vodde — resolves scaling by de-scaling: keep Scrum whole, just add the minimum extra machinery needed to run it with several teams working on the same product.

Two configurations — very different sizes

  • LeSS — up to about eight teams (≈50 people). One Product Owner. One Product Backlog. One Sprint. One shippable product Increment per Sprint. One shared Definition of Done.
  • LeSS Huge — many more teams (up to thousands of people on one product). Still one PO + one Product Backlog, but the backlog is split into Requirement Areas, each with an Area PO and an Area Product Backlog that draws from the one whole-product backlog. Requirement Areas are customer-visible business slices — not technical layers.

The LeSS shape

flowchart TB
  PO[One Product Owner]
  PB[(One Product Backlog)]

  subgraph Sprint["One Sprint (all teams synchronized)"]
    T1[Team 1<br/>feature team]
    T2[Team 2<br/>feature team]
    T3[Team 3<br/>feature team]
    T4[Team ...<br/>feature team]
  end

  DoD[One shared<br/>Definition of Done]
  INC[One shippable<br/>product Increment]

  PO --> PB
  PB --> T1 & T2 & T3 & T4
  T1 & T2 & T3 & T4 --> INC
  DoD -.-> T1 & T2 & T3 & T4

  classDef nodest fill:#f4f4f4,stroke:#444,color:#000
  class PO,PB,DoD,INC,T1,T2,T3,T4 nodest

Feature teams vs component teams — the AI programme tension

LeSS is emphatic: feature teams (cross-functional, own a slice of customer value end to end) are the default; component teams (own a technical layer) are an anti-pattern to be minimized.

Applied to an AI programme this is one of the most consequential org-design decisions you'll make:

  • Feature-team-first pattern — each team owns a model end-to-end: data pipeline in, training + eval, serving, integration into the consuming product. Cross-team specialists (data engineers, MLOps engineers) travel between teams via the LeSS "travelers" and Community-of-Practice mechanisms. Optimizes for whole-product agility; cost is that each team needs enough ML depth internally.
  • Component-team pattern (LeSS anti-pattern) — a "feature store team," a "training platform team," a "serving team," each owning a technical layer. Every model feature spans all three teams via handoffs. Common in AI orgs; produces the cross-team dependency tangles LeSS is built to prevent.
  • LeSS Huge Requirement Areas for AI — the correct dimension to split by is customer-visible AI product area (recommendations vs pricing vs fraud) not technical layer. Shared ML platform capabilities become a support area the whole product depends on, not a Requirement Area of its own.

One Product Backlog — the governance win

In LeSS, responsible-AI work (fairness audits, drift monitoring, guardrails hardening) sits on the same backlog as feature work, prioritized by the same PO against the same value hypothesis. There is no "governance backlog" that quietly slips behind features; the trade-off is made in the open at every Sprint boundary.

That single-backlog discipline is LeSS's most valuable property for AI — and the property SAFe deliberately doesn't have.

SAFe vs LeSS — side by side

Dimension SAFe LeSS
Kind of thing Prescriptive framework: roles, ceremonies, artefacts, configurations Minimal extension of Scrum: as little added as possible
Prescriptive or descriptive Prescriptive at every level (roles, cadences, artefacts) Prescriptive that nothing extra be added beyond named exceptions
Size sweet spot Hundreds to thousands (Portfolio + Large Solution SAFe) Up to ~50 in LeSS; thousands in LeSS Huge (still one backlog, one Sprint)
Number of Product Owners One PO per team + Product Manager per ART + Solution Mgmt per Solution Train One PO (or Area POs in LeSS Huge) for the whole product
Backlogs Team + Program + Solution + Portfolio backlogs (with WSJF prioritization crossing levels) One Product Backlog (or one + Area Product Backlogs in LeSS Huge)
Cadence 2-week iterations inside an 8–12 week PI (Program Increment) One Sprint across all teams (same length, same start/end)
Explicit platform-team support Yes — a platform ART is a common SAFe configuration No — platform capability is embedded in feature teams; LeSS explicitly distrusts component teams
Governance style Portfolio Lean budgets + Portfolio Kanban + WSJF The Product Owner's Sprint-boundary prioritization
Regulatory / audit fit Very strong — SAFe roles + PI Planning artefacts + Portfolio review map to audit expectations Modest — one backlog + one PO is defensible but the artefact set is smaller
Cost of running High — role weight, ceremony load, cross-level coordination Low — but very demanding on the culture (real feature teams, real cross-training)
AI-program fit Enterprise AI at scale (many products, many teams, regulated) Product-focused AI orgs where the whole-product agility is more valuable than portfolio governance
Compose with plan-driven governance SAFe inside a PRINCE2 stage / PMBOK life-cycle is a common hybrid (see Project management) LeSS inside a PRINCE2 stage is possible but has more friction (LeSS's minimalism resists ceremony import)

Choosing between them for an AI programme

  • Pick SAFe when you have several product ARTs and a shared ML platform ART, a portfolio to fund, a regulated environment that expects portfolio-level artefacts, and the org can absorb the ceremony weight. The AI Platform ART pattern only really fits SAFe.
  • Pick LeSS when you have one AI product (however big), a single PO who can hold the whole backlog, and an org culture that will actually build feature teams (not component teams wearing "feature team" hats). The single-backlog governance discipline is LeSS's superpower.
  • Hybrid with a plan-driven governance shell — either framework composes with a PRINCE2 / PMBOK outer shell (see Project management): agile delivery inside a plan-driven stage, with the Continued-Business- Justification decision at the stage boundary. Enterprise AI programmes almost always end up here.

Ties into the rest of Chiron

  • The SAFe "Continuous Delivery Pipeline" extended with the CT (Continuous Training) stage is provisioned by the Observability chapter (drift signals trigger retraining) and the Guardrails chapter (guardrail policy versions promoted alongside model versions).
  • WSJF prioritization at Portfolio SAFe folds Cost numbers (per-decision unit economics) into the cost-of-delay calculation.
  • The AI Platform ART in SAFe / the shared platform team in LeSS is the architectural target of the ADM Phase D work on the Architecture method page.
  • The feature-team / component-team choice made here directly shapes the AI Architect role's remit on the Roles & archetypes page.

Where the pillar goes next

  • Roles & archetypes — the capstone matrix of role × method × archetype, weaving the four preceding pages together with the AI Architect worked in depth inside the AI Transformation archetype.