The maturity assessment is finished. Forty-one capabilities, each scored one to five, rendered as a heat map in institutional colours. The consultants presented it well. The steering group nodded. The roadmap says most capabilities should reach Level 3 within three years.
You're the one who has to turn it into decisions. And there's a question the heat map can't answer: which of these scores is actually a problem?
The Employer Partnership Management capability scored a two. So did Timetabling. On the map, the two scores are identical. One might be the capability maturity gap that decides whether your strategy survives. The other might be exactly where that capability's maturity should sit: a deliberate two, in a capability where building maturity any higher would cost more than it returns. The heat map colours them the same. The maturity roadmap treats them the same. The uplift budget will be spread across both as though increasing maturity were a virtue in itself.
That's not a flaw in the scoring. The assessment measured how mature each capability is. Nothing measured how mature each one needs to be, and that second number can't come from the capability itself. It comes from what the organisation is trying to achieve: its strategy.
Capability maturity is the discipline of measuring how well an organisation performs each of its business capabilities, on a staged scale from ad hoc to optimised, and deliberately improving the ones that matter most. The discipline fails at a predictable point: deciding which capabilities matter most, and how mature each one actually needs to be.
Keep the map separate from the numbers. A capability map answers "what are we able to do?": it's the shared inventory, and it carries no scores at all. Every capability on it can then carry three numbers, each answering a different question from a different source. The assessment's score answers "how mature is this capability today?": call that number current. The strategy answers "how mature does it need to be?": that's required. And outcome evidence answers "how much of the claimed maturity is real?": that's validated.
The Three Numbers: current, required, validated. Your assessment delivered the first one, forty-one times, and the first one is the only one that can't change a decision. The other two are where the money moves: they're what tells a crisis two from a deliberate two, what turns a heat map you present into a budget you reallocate, and what gives you an answer the next time the steering group asks why these capabilities, at this cost.
Same current, opposite finding. The assessment delivers the first number forty-one times; the finding appears only once required is drawn beside it.
A capability maturity score without a strategically derived target is a number, not a finding.
And here is where most capability maturity roadmaps quietly fail. If the leadership team has never made its strategic choices explicit, where the organisation will compete and how it intends to win, then there is nothing to derive the targets from. The only reference left is the scoring model itself, so the roadmap points every capability at a higher level and calls that a plan. The organisations getting value from maturity work aren't scoring more rigorously than you are. They're setting targets differently, and you can spot the difference just by reading a roadmap, before a dollar of capability investment is spent. Targets derived from strategic choices naturally come out uneven: some capabilities pushed hard, others deliberately left alone. Targets defaulted from the scoring model come out uniform. If the roadmap in front of you says "everything to Level 3", you already know which kind you're holding.
The Maturity Theatre Diagnostic
Climbing maturity levels rigorously toward targets that were never anchored in a strategic choice is quite simply Maturity Theatre. The rubric is defensible, the uplift program is funded, governed, and reporting progress, and nobody chose the ladder any of it is climbing: nobody can name the strategic choice that requires these capabilities to reach these levels. It looks like diligence from the outside, which is exactly why it survives.
If you've read the OKR pillar or the benefits realization pillar, you'll recognise the shape. This is the same anchoring failure that produces OKR Theatre and Measurement Theatre, this time at the capability layer: rigorous apparatus, missing anchor.
Related reading: The Phase Every Management System Skips. All three failures reduce to one skipped phase: Define, the strategic choice none of these tools will make for you.
Five questions will show you whether your program is doing maturity work or Maturity Theatre. You can run them in under ten minutes.
1. Can you name the strategic choice that sets the target level for each capability you're investing in?
Take the three biggest lines in the capability maturity uplift roadmap. For each, complete the sentence: "this capability must reach this maturity level because we chose to win in this specific way". If the honest completion is "because the maturity model has five levels and we're at a two", the target is a default, not a decision. Uniform targets fail this test automatically: every capability to the same level by the same date is a target nobody derived.
2. Does your assessment score performance or deployment?
"The CRM is implemented and staff are trained" is deployment. "Enquiry-to-enrolment conversion improved, and we know by how much" is performance. If the scoring rubric awards levels for systems existing, processes being documented, and training being completed, rather than for outcomes moving, a capability can reach Level 4 while performing no better than it did at Level 2. Healthcare's maturity practitioners learned this the hard way and gave the failure a name; the cases below cover what it costs.
3. Is any capability deliberately targeted to stay immature?
A genuine strategy produces capabilities that are deliberately held low, for one of two reasons. Some are commodities: good enough is genuinely enough, and every dollar spent maturing them past that point is a dollar taken from the capabilities that carry the strategy. Others are exploratory: the value comes from experimentation, and adding process weight would slow the very thing that makes them valuable. Either way, the decision is explicit, and someone can say which reason applies. If every capability's target is at or above its current level, no such decision was ever made, and the roadmap is treating maturity as universally good. That's the same as never having asked what each capability is for.
4. Did the last maturity assessment change an investment decision?
Name one. A funding line that moved. A duplication that was consolidated. An uplift that was cancelled because the capability was already mature enough for what the strategy asks of it. If the heat map's only output was the heat map (presented, admired, filed), the assessment produced a deck, not a decision. The map isn't the value. The reallocation is.
5. Is your business case built on your own baseline, or on someone else's ROI multiple?
The point of this question: a business case built on borrowed numbers can never be proven right or wrong, and a case that can't be proven wrong protects the program instead of the organisation. The classic figures trace to a thirteen-organisation software study from 1994, drawn from a literature that has never once published a loss; the borrowed number describes someone else's best case and predicts nothing about yours. The only evidence that can ever prove your program worked is your own baseline: current performance, measured, per capability, before the program starts.
Each question probes one of the Three Numbers or a difference between them. If three or more don't have a clean answer in your context, the program is doing Maturity Theatre, and now you can say precisely what that means: it's a one-number program. The scoring is probably sound; what's missing is the other two numbers, and the fix isn't a better rubric. It's deriving targets from strategic choices before the next capability investment is committed.
The Maturity Contour
A flat maturity roadmap isn't a plan. It's written evidence that no strategic choices were made.
The fix has a shape. The right maturity profile across a capability map is deliberately uneven, contoured by the organisation's strategic choices.
One thing needs pinning down first: what the scale actually measures, because everything that follows depends on it. A maturity level measures how institutionalised a capability's operation is: how repeatable, instrumented, and embedded its people, process, technology, and information are, from ad hoc effort at the bottom of the scale to measured, continuously improving operation at the top. Performance evidence is what validates a maturity level claim: a capability hasn't genuinely reached Level 4 unless the outcomes it exists to produce are measurably better than they were at Level 2, and deployment evidence validates nothing.
Maturity is a relation, not a property. The scale gives each capability its Three Numbers (where required comes from is shown two sections ahead), and the findings live in the differences. Current below required is a strategic gap, the only kind that deserves the word "crisis". Current above required is over-investment: money spent climbing a wall the strategy never asked to be climbed. Claimed current above validated is the Infrastructure Trap, deployment dressed as capability. And current equal to required, on purpose, is the deliberate two: keep the operation light, not the outcomes poor. That's the reading discipline to hold from here on, whatever rubric your assessment used.
With the numbers named, two pieces need to be in place before the contour can be drawn: a shared map, and a source for the targets.
Map before score
Mapping and scoring are complementary layers, not competing frameworks. A map with no maturity overlay is a static picture; a maturity assessment with no shared map can't be compared, benchmarked, or acted on consistently. The field itself has split three ways over this distinction: staged maturity models that score levels but can't say which level you need, capability reference models that map what you do but stay deliberately silent on how well, and outcome-based capability models that reject levels entirely. Each camp holds a piece. The contour uses all three: the shared map, the scale, and the rule that only performance evidence validates a level.
Higher education is the demonstration case because the sector already solved the hard half. The Higher Education Reference Models (HERM), curated by CAUDIT and used by more than a thousand institutions worldwide, give the sector a standard Business Capability Model, and HERM is deliberately maturity-agnostic: it ships the map, not the rubric, so the shared taxonomy stays stable while each institution assesses itself against its own strategy. Banking has BIAN, telecommunications eTOM; the shared-map pattern generalises. What no sector can share is the target levels, because targets are derivatives of strategy, and strategy is the one thing peer institutions don't have in common.
The shared map also unlocks like-for-like benchmarking, and benchmarking's job needs stating precisely, because it is not target-setting. A benchmark is a ruler, not a decider: it can tell you where you sit relative to peers, never where you need to sit, because that relation is the strategy's to choose. So when the contour puts a strategy-carrying capability above the sector norm, the target is still strategy-derived. The strategy did the deciding; the benchmark just does the measuring. (HERM's Data Reference Model does the same job for the information layer; the higher education data architecture pillar covers that half of the story.)
Where targets come from
A capability map becomes a decision tool when each capability is assessed on three dimensions: strategic importance, maturity, and adaptability. The third returns at the end, because it turns out to measure something bigger than a capability. The first is the anchor, and it can't be settled by interviewing capability owners, because every owner believes their capability is strategic. It has to be derived in the room where the strategic choices live.
The instrument that derives it is the Strategic Choice Cascade: winning aspiration, where to play, how to win, required capabilities, management systems. The fourth choice is where target maturity levels are born. An organisation that has answered where to play and how to win has already determined, whether it has written it down or not, which capabilities must be distinctive, which need only be competent, and which should be deliberately minimal.
One boundary condition before the cascade takes over: some floors aren't the strategy's to set. Regulators fix minimum maturity for capabilities like cybersecurity, privacy, and research ethics regardless of how the organisation chooses to win. Those floors don't break the rule that targets come from strategy: choosing to operate in a regulated space is a where-to-play choice, and the floor arrives with it. The contour is drawn above the floors, never instead of them.
An organisation whose strategy is a list of priorities rather than a set of integrated choices has no way to derive strategic importance. Its maturity targets will be set by the only forces left: the maturity model's ladder, the loudest capability owner, and the consultant's benchmark.
A target level, derived
Watch the derivation happen once, because the whole discipline is in it. What it answers is the Target-Level Question, the question every heat map raises and no maturity model can answer: how mature does this capability need to be, and which strategic choice says so?
An institution's how-to-win: the best-supported first-year experience in the region. The capability that claim lands on: student support, and specifically its early-alert dimension, the ability to see a struggling student and intervene before withdrawal. The strategy's claim is comparative (best-supported in the region), which makes the sector norm the thing to beat. Peer institutions run early alert at a defined-but-unmeasured level: the process exists, interventions happen, nobody instruments the outcomes. Call that the benchmark three.
A three can't carry the claim, because matching peers wins nothing. A five, a fully optimising operation redesigning itself term over term, is more than the claim requires and would consume investment the cascade has committed elsewhere. The target is a four: instrumented, measured intervention, validated by outcome evidence, in this case the retention of the students the system flags. Hold that four as provisional: a per-capability target is a first pass, and it becomes final only when the value-stream walk further down confirms the handoffs it depends on.
Now complete the diagnostic's sentence: this capability must reach Level 4 because we chose to win on first-year support, and winning means being measurably better than a sector that operates at three. That sentence, repeated for the handful of capabilities the strategy actually depends on, is the Define-phase output your assessment was always missing.
Free worksheet: The Maturity Contour Worksheet runs this derivation across a whole map — the Three Numbers per capability, the strike, floor, peak, plateau and valley walk, and the choosing sentence every row has to survive.
Peaks, plateaus, and valleys
Run that derivation across the map and the contour appears: the Three Numbers drawn for every capability, with the differences made visible. It has three features.
Peaks are the small number of capabilities where the organisation's how-to-win lives; the early-alert four just derived is one. These must sit above the sector norm because the strategy's claim is comparative, and underinvestment here is an existential gap: a three at a peak is a crisis regardless of what the benchmark says.
Plateaus are the table-stakes capabilities every peer runs. Payroll, procurement, timetabling need to be reliably boring: standardised, efficient, probably indistinguishable from any peer's. The heat map's most valuable output here is the duplication finding. Three faculties running the same capability three different ways is a consolidation decision waiting to be made. And investing a plateau past "efficient and reliable" is spending distinctiveness money on a capability that can't return distinctiveness.
Valleys are the deliberately loose capabilities: new-program incubation, exploratory research directions, novel partnership models. Stage gates and maturity uplift here would slow the experimentation that produces the next peak. A valley is a deliberate two at map scale: a chosen low-maturity state, with a rationale, a boundary, and a review rhythm. When an experiment becomes an operation, it graduates onto the contour. Neglect is different: neglect is the absence of a decision. A neglected two and a deliberate two look identical on a heat map, and the difference is whether anyone can produce the sentence that chose it.
The walk has two more possible verdicts, and neither is a target level. Strike: a reference model like HERM catalogues everything an institution could do, and your where-to-play choices select the subset you actually do; an institution with no clinical programs doesn't hold clinical trials management at a deliberate two, it strikes it, with the striking sentence on record. Source: some capabilities you need performed but don't need to own, and the commodity plateau is the natural candidate. The target doesn't disappear when you source a capability out; it moves to managing the provider. The one thing that can never be sourced out is a peak: if a vendor can supply your how-to-win, it was never a how-to-win.
The flat roadmap you were given, against the contour your strategy implies. Red marks where the flat plan under-invests the strategy; the plateaus, valley, and the struck and sourced capabilities are all deliberate choices.
Higher maturity is not automatically better. Maturity discipline pays where predictability pays: stable, repeatable, high-volume capabilities. It costs where discovery pays. A capability maturity model can't know which is which. A strategy can.
One caution as you draw it: the contour prescribes altitude, not remedy. Reaching a peak's four is rarely a matter of adding process alone, because a capability is four elements, and the gap can live in any of them: nobody with employer-facing skill is a people gap, no shared definition of a partnership across faculties is an information gap, a system that can't carry the load is a technology gap. The contour says how high. The gap classification in the next section says what kind of climbing.
The second pass: walk the value stream
The cascade walk derives targets capability by capability, and that has a blind spot: it can't see a broken handoff between two individually adequate capabilities. So the contour needs a second pass: take each value stream that carries the how-to-win, walk it end to end, and check every capability it crosses, because a nominal plateau sitting on a strategy-carrying stream inherits requirements the per-capability walk will never produce. Take timetabling, the example of reliably boring from a moment ago: at an institution whose how-to-win is work-integrated learning, it sits directly on the stream that carries the strategy, because placements live or die on scheduling. The plateau target survives the second pass only if the handoffs do. This is also where the Deliver question gets its method: a service blueprint of the strategy-carrying stream is how you test whether maturity survives contact with the person it exists for.
Capability Maturity and the Four Ares
The contour isn't a fifth phase bolted onto the Design4 framework. Each phase of the cycle already asks a maturity question:
| Design4 Phase | Four Ares Question | Maturity question | What a rigorous answer requires |
|---|---|---|---|
| Discover | Are we getting the benefits? | Is the capability delivering outcomes to the stakeholders it exists to serve? | Performance evidence connected to purpose: conversion rates, retention, outcomes moving. Not adoption counts, not deployment milestones. |
| Define | Are we doing the right things? | Which capabilities must be how mature, and which stay deliberately loose? | Target levels traceable to specific choices in the cascade. The contour, in writing, with the valleys as explicit as the peaks. |
| Develop | Are we doing them the right way? | Is the capability an integrated system of people, process, technology, and information at the required level? | Assessment across all four elements, with each gap classified before it's remedied. |
| Deliver | Are we getting them done well? | Does the capability's maturity survive to the stakeholder experience? | Service-blueprint evidence from the value-stream second pass. A capability can score Level 4 internally while the stakeholder's journey across its handoffs fails. |
Each phase already asks a maturity question, and asking it catches a specific failure: the Infrastructure Trap at Discover, flat targets at Define, misclassified gaps at Develop, broken handoffs at Deliver.
Two implications are worth pulling out of the table.
A low maturity score is a symptom with three possible diseases. The Design4 pillar's gap taxonomy distinguishes a maturity gap (you have the capability, it isn't good enough), a resource gap (you don't have it), and an integration gap (the pieces exist but don't connect); each needs a different remedy, and a score alone can't tell you which you're looking at. Treating a resource gap with process improvement, or an integration gap with a technology purchase, is how organisations spend years addressing the wrong problem. The score tells you where to look. The classification tells you what to do.
The cycle is what makes the ladder safe. A capability maturity model without Define produces uniform climbing. Without Discover, it produces the Infrastructure Trap. Without Deliver, it produces capabilities that are mature on paper and broken at the handoffs. The model is a fine instrument inside the cycle and a hazardous one outside it, and no maturity model has ever recommended its own non-use.
Nobody Publishes a Failed Maturity Program
One warning before the case studies, because the business case that bought your heat map probably cited a return multiple. In three decades of published research on maturity program returns, no study has ever documented a loss; failed programs are real, but they don't get written up, so the literature is a record of what the survivors returned. The classic benchmarks trace to thirteen software organisations in 1994, and the stronger modern evidence (healthcare's digital-maturity outcome studies) is honestly correlational. The full argument, and how to write a maturity business case the evidence actually permits, is here.
The short version: expect a plausible range rather than a headline multiple, budget for the J-curve because new process dips productivity before it raises it, and gather your own baseline before the program starts, peaks first. With a baseline, the program has to earn its next funding round. Without one, any improvement is indistinguishable from noise and any failure can be explained away.
What This Looks Like in Practice
The heat map that never left the deck
A university commissions a maturity assessment on its HERM-based capability map. Forty-one capabilities scored. The findings are genuinely good: rigorous rubric, wide interviews, defensible scores. The roadmap recommends bringing every capability below Level 3 up to Level 3 within three years, prioritised by score, lowest first.
Eighteen months later, two uplift projects are running, both in back-office capabilities that scored lowest. The heat map hasn't been refreshed. And the strategy, which commits the institution to distinctive work-integrated learning, isn't mentioned anywhere in the uplift portfolio. The capability that carries that strategy, employer partnership management, scored a two and sits eleventh in the queue. Behind records management. A current of two, and a required nobody ever derived.
Somewhere in the third year of that queue is a student who chose this institution for the placement it promised her, and who will graduate without one, because the capability that finds placements ran on two overworked coordinators and a spreadsheet while the uplift money went to records management. She is what a misdrawn contour costs. Nothing on the heat map shows her.
The fix runs through Define, not through re-scoring, and it needs a sponsor before it needs a session. The provost, whose strategy the roadmap is quietly failing, convenes the working session and owns what it decides. The session walks the Strategic Choice Cascade and extracts the four capabilities the how-to-win actually depends on; their targets are set from the strategy, not the model, and two need to exceed the sector norm. Then the session walks the rest of the map with the opposite question: which capabilities need less than the roadmap assumed? Three plateaus are duplicated across faculties, surfacing a consolidation decision nobody had framed, and the innovation-facing capabilities become valleys, with protect-from-process decisions and review dates.
The queue re-sorts from "score ascending" to "distance from strategically required level, weighted by importance". Employer partnership management moves from eleventh to first. Two queued uplifts are cancelled because the capabilities are already sufficient for what the strategy asks of them, which nobody had previously had grounds to say; on the new roadmap they become deliberate twos, current equal to required on purpose, with the choosing sentence on record. The politics are real: the owners of the cancelled uplifts had been promised investment, and the session where the contour overruled the queue was not a comfortable one.
The mechanics matter as much as the logic. A promised uplift is a governance commitment, so the re-sorted queue goes back to the steering group that approved the original roadmap, on the provost's motion, because only the body that made a commitment can unmake it. The contour gets a named owner and a standing slot in the annual planning cycle, and the assessment finally produces what it was bought for. A reallocation.
The program that deployed its way past the capability
The National Programme for IT (NPfIT), launched by the English NHS in 2002, is the cleanest public record of deployment mistaken for capability at national scale.
The goal was integrated electronic care records, procured centrally as the largest civilian IT program in the world at the time, on the assumption that the capability would follow the technology: roll out the systems, and clinicians would soon be using shared records to deliver safer, better-coordinated care. It didn't happen, because deployment was the only layer being governed. The people element (clinician engagement) and the process element (enormous local variation in clinical workflow) were treated as change-management footnotes, and progress was reported in contracts signed and systems delivered. The actual capability question, whether a clinician in one setting could reliably see and use a record created in another, went unasked at governance; that unasked question is the information element, the dimension the benefits realization pillar calls the last to be tested and the first to fail in production. The program was dismantled in 2011, condemned by the Public Accounts Committee as "one of the worst and most expensive contracting fiascos in the history of the public sector", its costs, forecast at £9.8 billion and still rising when the committee reported, never shown to justify the benefits.
Readers of the benefits realization pillar will recognise the rhyme with Phoenix: both programs had green dashboards, and both dashboards were answering the wrong layer's question. Phoenix wrote a strategic cheque against an unverified capability account; NPfIT is the same failure one layer down, a capability case built as a technology case, one element of a four-element system mistaken for the whole. In the Three Numbers: a claimed current with validated at zero, discovered on launch day.
Healthcare, with its eyes open
Healthcare has taken maturity discipline further than any other sector, and it works as the positive case precisely because it also supplies the cautions.
The HIMSS staged models, led by the Electronic Medical Record Adoption Model (EMRAM), score a hospital from Stage 0 to 7 against a vendor-neutral roadmap, backed by outcome evidence across a thousand-plus hospitals. And funders are starting to attach money to it: Germany ties hospital modernisation funding directly to demonstrated digital-maturity progress, and NHS England runs a national digital maturity assessment whose rankings increasingly steer central investment. Maturity has graduated from assessment exercise to strategic lever, and higher education's regulators and funders could plausibly move the same way. A forward-looking reason to build the discipline before it's imposed.
The caution comes from inside the same sector. Healthcare's own critics coined the term Infrastructure Trap for maturity models that equate maturity with technology adoption and under-weight the human and organisational capabilities that determine results. In house terms, the Infrastructure Trap is Measurement Theatre's Develop-phase twin: deployment evidence promoted to capability evidence without anything actually performing better. NPfIT is what the trap costs at scale. The lesson isn't "adopt EMRAM for universities"; it's the pattern with the caution built in: shared map, outcome-evidenced scoring, targets derived from strategy, and performance, never deployment, as the thing scored.
The regional health authority case in the Design4 pillar shows the whole discipline compounding: a value stream located where capability handoffs broke (an integration gap no maturity score would have caught), the Four Ares re-anchored investment to patient outcomes, and each governance cycle redrew the picture faster than the last. Maturity gave that organisation the vocabulary for what to fix; the cycle told it where fixing paid.
The contour isn't a higher education device. A bank working on BIAN's map finds its payments plateau duplicated across three product lines (a consolidation decision waiting to be framed), its how-to-win living in two data capabilities that must exceed the peer norm, and a deliberate valley around an embedded-finance experiment that heavier process would kill. Different map, same walk.
The Question Above the Ladder
One question has been sitting underneath everything above, and no maturity model asks it: how mature is the process that decides how mature your capabilities need to be?
Apply the scale to itself. An organisation whose contour was drawn once, by consultants, and filed has an ad hoc deciding process: a Level 1 organisation, whatever its capability scores say. An organisation where the derivation, the validation, and the redraw run on a governance calendar, with a named owner and decisions on record, is more mature than its peers even if half its map sits at deliberate twos, because maturity was never about the height of the bars. It's about whether anyone governs the relation between them. That's what adaptability, the third assessment dimension, actually measures at map scale. It's also what the AI transition, with its epidemic of perpetual pilots, is grading right now. The full argument deserves its own page: How Mature Is the Process That Decides?
The short version: the level was never the product. The product is the durable capacity to keep converting whatever arrives next, that capacity is the account your strategy keeps drawing on, and the contour, redrawn every cycle, is how the balance keeps growing.
Go Deeper
You now have the reading discipline: the Three Numbers, where required comes from, the contour's features, and the question above the ladder. What you don't yet have is fluency with the map itself. The working session in the case study moves fast because someone in the room can hold HERM's capability model and translate between its taxonomy and each executive's language; without that fluency, the session produces a debate about definitions instead of a contour.
Building the Common Language develops exactly that: reference model literacy, with HERM as the working example, and the discipline of making a shared taxonomy speak in each audience's terms. Closing the Strategy-Execution Gap covers the upstream half of the discipline: the Strategic Choice Cascade that derives the targets, the heat map and gap classification, and the governance rhythm that keeps the contour current. Deriving the contour and holding it are different skills, and the second one gets tested in the first funding round after the re-sort, in a room where every incentive pulls toward restoring the queue; that holding is what the course's governance work is for.
If the diagnostic told you your roadmap is flat, the cascade course is where the targets come from. If the targets exist but the map keeps dissolving into vocabulary arguments, the reference model course is the missing skill, and Core Means Two Things on Your Capability Map names the specific argument before it starts.
Frequently Asked Questions
Isn't this just CMMI?
CMMI is a scoring discipline. The contour is about where the targets come from, which is a question CMMI is structurally silent on, because a cross-industry model can't know your strategy. Notably, CMMI's current release has itself been reframed around independently verified performance outcomes rather than compliance: the same correction the contour makes. Use CMMI-style rubrics happily, inside a frame that derives each capability's target from a strategic choice and permits the answer "this capability should stay at Level 2, on purpose".
Aren't maturity models fundamentally broken?
That argument, made most influentially by Accelerate and amplified across the DevOps and agile communities, holds that staged maturity models imply a false finish line, impose activity checklists instead of measuring outcomes, and ignore context. Every part of the critique is right, and every part of it describes Maturity Theatre: an unanchored ladder, climbed for its own sake, scored on deployment. The contour concedes the whole indictment and keeps the one thing the critique discards: a shared scale. Targets still need to be communicated to boards, compared across units, and governed over years, and raw outcome metrics can't do that job alone. What's broken isn't the scale. It's the assumption that the top of it is everyone's destination.
We already have a HERM capability map. What do we actually do next?
Overlay, don't rebuild. First strike the capabilities your where-to-play choices exclude: HERM catalogues everything an institution could do, and your map is the subset you actually do. Then assess each remaining capability on three dimensions: strategic importance (derived from the choice cascade, not scored by interview), maturity (performance across people, process, technology, and information, not deployment), and adaptability. Assess per unit, not institution-wide. A single university-wide number hides the fact that one faculty runs a capability at Level 4 while another runs the same capability ad hoc, and that variation is usually the most actionable finding on the map, because spreading the mature practice is cheaper than any uplift program.
The board wants one maturity number. What do I give them?
Give them the contour, not a number. A single number averages peaks against valleys and reports noise. The board-ready view: the handful of strategy-carrying capabilities with target versus current and trajectory, the duplication and consolidation findings, the valleys as explicit choices, and one or two outcome indicators showing whether rising maturity is moving anything that matters. "Our average maturity is 2.8" tells the board nothing. "The four capabilities our strategy depends on are here, need to be here, and here's the evidence the gap is closing" is a governance conversation.
How does this connect to the Benefits Stack?
Directly: capability maturity is the Layer 2 discipline the benefits realization pillar depends on. That pillar showed that a strategic benefit is a cheque drawn on a capability account; maturity assessment is how the balance gets verified before the cheque is written. The two disciplines meet in the middle of the stack.
Which KPIs should I actually track for each capability?
The capability map answers this before you pick a single metric. A capability with no value stream or strategic choice behind it has no business owning a KPI at all, and that is exactly how organisations end up with two hundred measures and nothing that is actually "key." For each capability on the contour, ask three questions before adding a metric: what strategic outcome does this inform, what value stream does it sit on, and who acts when it changes. If you cannot answer all three, the KPI has not earned its place. This is the same discipline the Business Architecture Guild's BIZBOK and the Open Group's Business Architect competency model both formalise: capabilities and value streams are meant to be the structure KPIs get derived from, not an afterthought layered on top of whatever was easy to measure. It also solves the ownership problem that sinks most KPI programs: a metric like "student satisfaction" has no natural owner because it was never derived from a capability in the first place. Trace it back to the value stream and the capability that actually produces it, and the owner is the capability owner, not a debate.
The Deliberate Two
The heat map on your wall right now probably contains both kinds of two. One is a crisis wearing the same colour as a choice.
The practitioners who get value from capability maturity aren't the ones with better rubrics. They're the ones who walked the strategy into the room, derived the targets before defending the scores, and then held the contour in the meeting where it cancelled work someone had been promised. Seeing the difference between the two twos is the skill. Holding it, in front of the owner whose uplift just became a deliberate two, is the practice.
You'll know the work has landed the first time you point at a two in a steering committee and the room understands it's there on purpose. That moment is the visible test. The invisible asset behind it is the process that produced the sentence: an organisation that can now decide, on a calendar, with evidence, how mature anything needs to be. That's not a scoring skill. It's an anchoring skill, and it's the one the next assessment, the next funding round, and the next technology wave will all grade.
Start Monday morning. Pull up the uplift roadmap, find the flattest stretch of targets on it, and ask who chose them. If nobody can produce the sentence, you've found your entry point.
Continue Learning
Sibling pillars: Design4 sets out the four-phase cycle this pillar's maturity question attaches to at every stage. Benefits realization is the Layer 2 discipline this contour feeds, the account a strategic benefit draws on.
In the cluster: How Mature Is the Process That Decides? takes the question above the ladder and runs it as its own argument. Core Means Two Things on Your Capability Map names the vocabulary fight before it starts.
If you're building a contour, The Maturity Contour Worksheet is the companion template.
Building the Common Language and Closing the Strategy-Execution Gap teach the two halves of holding a contour: reading the map fluently, and deriving and governing the targets on it.