Skip to content

Traditional iterative development

Plan-driven, iterative is a middle ground: not agile, not waterfall. You commit to an overall architecture and a phased rhythm, but you deliver in short iterations that retire risk. It is the delivery style regulated and high-assurance AI programs default to when a pure Scrum cadence would be too loose and a waterfall would be catastrophic.

Two frameworks dominate:

  • Rational Unified Process (RUP) — a process framework: four phases × nine disciplines, iterated within each phase, with the discipline effort visualized as the "hump chart."
  • Boehm's Spiral model — a risk-driven meta-model: every loop identifies and retires the biggest remaining risk before committing more resources.

Both predate modern agile; both are alive and well in AI programs where the top risk changes each iteration (data quality → model feasibility → integration → adoption). This page presents RUP as the primary framework and Spiral as the peer, as linked method tabs — pick RUP once and every tabbed block on the page follows.

Sources verified 2026-08-09

  • RUP (Rational Unified Process) — IBM/Rational; canonical overview at en.wikipedia.org/wiki/Rational_Unified_Process. Foundational IBM/Rational whitepaper: "Rational Unified Process — Best Practices for Software Development Teams" (Rational Software Corporation, 1998, TP026B).
  • Spiral model — Barry W. Boehm, "A Spiral Model of Software Development and Enhancement," IEEE Computer 21(5), May 1988, pp. 61–72 (DOI 10.1109/2.59). Wikipedia overview: en.wikipedia.org/wiki/Spiral_model.

The AI-program lens

The AI programs that best fit iterative-but-plan-driven delivery share three features:

  • Feasibility is genuinely uncertain. You don't yet know whether a model can hit the metric on the data you actually have. That uncertainty must be retired before you commit to full Construction.
  • The risk hierarchy is stable. Data risk almost always dominates, then model risk, then integration/latency, then adoption. A framework that sequences work by biggest remaining risk fits the way AI programs actually fail.
  • Governance requires a defensible artefact trail. Regulators and internal audit expect signed-off phase deliverables (an architecture, an eval report, a validation plan). Sprint-only delivery struggles to produce those cleanly; RUP and Spiral produce them by construction.

Both frameworks below are used through this lens.

The two frameworks

RUP is a 2-D process framework: four sequential phases on one axis, nine disciplines on the other, and iterations inside each phase. Effort in each discipline rises and falls across the phases — the shape is the famous hump chart.

The four phases (canonical order + objective)

Phase Purpose AI-program lens — the risk this phase must retire
Inception Scope the system enough to validate the business case + initial costing. Is this problem worth solving with a model? Value hypothesis, decision-point, KPI-to-move; data-availability first-pass; go/no-go on the whole program.
Elaboration Mitigate the architecturally-significant risks. Produce a validated architecture baseline. Model feasibility spike: a PoC that proves the model can hit the KPI on the real eval set with the real data (or the closest available). If the PoC misses, kill Elaboration — don't roll into Construction hoping training will save it. Also: pick the accelerator strategy, the serving pattern, and the ground-truth-labeling operating model.
Construction Build the software system to a released beta. The train-eval loop runs at full cadence: data prep, feature engineering, training runs, systematic evaluation with regression gates. The inference service, the guardrails layer, and the human-review UI reach beta.
Transition Move the system into production for real users. Shadow mode → canary → progressive rollout with a live rollback trigger tied to a quality metric. Ground-truth backfill during shadow. On-call rotations for model serving. Documentation for the operators who will retrain the model when it drifts.

The nine disciplines

Six engineering + three supporting; each runs (with varying effort) across every phase:

Engineering disciplines Supporting disciplines
Business Modeling Configuration & Change Management
Requirements Project Management
Analysis & Design Environment
Implementation
Test
Deployment

The hump chart — AI-adjusted

In generic RUP the engineering disciplines "hump" at predictable points: Requirements peaks in Inception/early Elaboration; Analysis & Design peaks in Elaboration; Implementation and Test peak in Construction; Deployment peaks in Transition. AI programs modify the shape — the humps below reflect where effort actually lands on a well-run AI project.

xychart-beta
    title "RUP hump chart — AI-program effort by phase"
    x-axis ["Inception", "Elaboration", "Construction", "Transition"]
    y-axis "Relative effort" 0 --> 10
    line [4, 8, 2, 1]
    line [7, 6, 3, 1]
    line [2, 9, 4, 2]
    line [1, 6, 9, 3]
    line [1, 5, 8, 5]
    line [1, 2, 4, 9]

Reading the lines top to bottom in Elaboration (peak-order): Analysis & Design (peaks in Elaboration — the architecture baseline including the model-feasibility spike), Business Modeling (front-loaded — the value hypothesis and decision point), Requirements (front-loaded but persists — eval-metric thresholds and NFRs keep evolving), Implementation (peaks in Construction — the train-eval loop at cadence), Test (peaks late Construction / early Transition — regression eval and shadow-mode measurement), Deployment (peaks in Transition — canary, rollback wiring, MLOps handover).

Elaboration is the phase AI programs skimp on — and pay for later

The most common AI-program failure mode is racing from Inception into Construction because "we know the model works on a demo." Elaboration exists specifically to prove the model works on your production-shaped data, at production latency, inside your governance envelope. If you skip it, Construction becomes a training-loop-until-we-give-up.

Iterations inside phases

Each phase runs one or more iterations. In Elaboration, an iteration is typically a model spike: pick the biggest open modeling risk, build the smallest thing that would disprove feasibility, run it against the eval set, land the result in the architecture baseline. In Construction, an iteration is a train-eval cycle ending in a signed-off eval report.

The Spiral model is not a process framework; it's a risk-driven meta-model. Every cycle is one loop through four quadrants — Boehm's 1988 paper. Which cycles you actually run, and what happens inside them, is determined by the risks the previous cycle exposed.

One cycle — four quadrants

flowchart LR
  Q1[Q1. Determine objectives,<br/>alternatives, constraints]
  Q2[Q2. Evaluate alternatives;<br/>identify + resolve risks]
  Q3[Q3. Develop and verify<br/>next-level product]
  Q4[Q4. Plan the next<br/>phase / cycle]
  Q1 --> Q2 --> Q3 --> Q4 --> Q1

Each successive cycle sits farther from the origin — cumulative cost grows outward while risk shrinks. Progress is measured inward: how much risk you've retired, not how much code you've written.

An AI-program spiral — the risks in the order you retire them

Cycle Objective this loop Alternatives evaluated Risk this loop retires What you build & verify
1. Data risk Do we have the data we need, at the quality we need, that we're legally allowed to use? Buy vs. label; internal vs. synthetic; full population vs. representative sample Data availability + quality + rights — the risk that no amount of modeling will save you A data audit + a labeled eval set + a data-contract prototype
2. Model risk Can any model hit the KPI on this data? Classical ML vs. small LLM vs. hosted frontier LLM vs. hybrid; feature engineering vs. embeddings Model feasibility — can the accuracy/quality bar be met at all A model PoC + a regression eval harness + a decision on modeling family
3. Integration / latency risk Can the model be embedded in the target system inside its latency and throughput budget? Sync vs. async serving; batch vs. real-time inference; on-prem vs. hosted; caching vs. retrieval Inference performance under production shape — tail latency, cold-start, retry storms, cost per call A serving prototype + a load-tested reference deployment + observability plane
4. Adoption risk Will the humans who consume the model output actually change their decision because of it? Human-in-the-loop vs. autonomous; recommend vs. auto-apply; explain vs. score Behavioural adoption — the model that isn't trusted is not used A shadow-mode rollout + user-behaviour telemetry + an explanation UX PoC

Cycles beyond #4 are the ones a mature AI programme adds because its risks are different (data-freshness risk, regulatory risk, model-supply-chain risk, cost-blowout risk). The framework is honest about not being prescriptive — you list your program's top risks and take a loop per risk.

Where each cycle ends

Every loop closes with Q4 — Plan the next phase, and that plan includes the honest option STOP. Spiral's discipline is that a cycle whose risk didn't shrink is a signal to abandon that direction, not to iterate harder.

RUP vs Spiral — side by side

Dimension RUP Spiral
Kind of thing Process framework (phases × disciplines) Risk-management meta-model
Prescriptive or descriptive Prescriptive process, tailorable via role/phase configurations Descriptive: says think about risk each cycle, silent on how
Primary artefact The 2-D matrix (phases × disciplines) + the hump chart The spiral diagram + the four quadrants
Unit of work An iteration inside a phase One full loop through the four quadrants
What it forces you to do Follow the phase sequence, produce phase deliverables, iterate inside each phase List your top risks each cycle, evaluate alternatives, retire the biggest risk before committing more resources
What it doesn't tell you Which risks matter (Elaboration is labelled risk-mitigation but doesn't rank the risks) What activities to actually perform inside a cycle (RUP or agile fills that in)
AI-program fit The default rhythm for programs that need signed-off phase deliverables (regulated / enterprise) The default risk-decision layer on top of any process — including RUP itself
Adopt when You need a repeatable disciplined delivery cadence with governance artifacts You have several plausibly-fatal risks and need discipline about retiring them in order
They compose because The RUP Elaboration phase is a small spiral by construction: Boehm-style risk cycles running inside the RUP phase. Modern practice runs Spiral as the risk-ranking meta-loop around RUP or a scaled-agile cadence.

Contrast with agile — and where the hybrid lives

  • RUP is not agile. It shares agile's iterativeness, but the phase gates and the volume of prescribed artefacts are much heavier than a Scrum cadence. In practice most modern shops that say "we do RUP" run OpenUP or a slimmed-down configuration.
  • Spiral is not agile either. It's a risk-decision meta-model that pre-dates agile. A Scrum team that ends every sprint with "what's the biggest remaining risk, and did we shrink it?" is running a de-facto WinWin Spiral cadence — but that's an earned discipline, not a Scrum default.
  • Where the hybrid lives. Enterprise AI programs typically overlay a plan-driven governance shell (PRINCE2 / PMBOK stage gates — see the Project management page) on top of an agile delivery cadence inside the program teams, with Spiral risk-ranking at the portfolio review. That's how you satisfy audit AND ship in short cycles.

The scaled-agile version of the same trade-off is covered in the next page in this pillar (Scaled agile).

When to reach for which — and how to compose

  • Reach for RUP when the program has enough novelty to need Elaboration (a real feasibility spike), enough scope for Construction discipline (multiple disciplines running at cadence), and enough regulatory weight that the phase artefacts aren't overhead — they're the audit trail.
  • Reach for Spiral when the program's failure modes aren't execution — they're strategic. If you can build what you design but you're not sure the right thing to build is the model you have in mind, Spiral's risk cycling is the frame that keeps you from building the wrong thing efficiently.
  • Compose them by running Spiral cycles inside RUP's Elaboration phase; running Spiral again at the phase gate at end of Construction to decide whether Transition is even the right next move; and treating the model itself as a product to be retired if a later Spiral loop shows it doesn't move the KPI.

Ties into the rest of Chiron

  • Elaboration's feasibility-spike output feeds directly into the ADM Phase C (Information Systems) architecture baseline from the Architecture method page.
  • Transition's shadow/canary + rollback wiring is the pattern covered in Observability and the guardrail rollout logic in Guardrails.
  • Spiral Cycle 4 (adoption risk) sits on top of the Trust & attribution approval workflow — the model that adoption depends on is exactly the one whose writes must be attributably approved.

Where the pillar goes next

  • Project management — schedule, governance stages, and CPM through an AI-program lens (PMBOK + PRINCE2).
  • Scaled agile (SAFe + LeSS) — scaling AI delivery across many teams.
  • Roles & archetypes — the capstone matrix of role × method × archetype.