AI foundations

AI maturity models: what each stage actually looks like

The head of AI at a mid-size logistics group opens the board pack with a slide: fourteen pilots delivered this year, three approved for phase two. He does not mention that phase two, for the two pilots that got there last year, never actually happened. The data scientist who built them left in March. Her laptop, and whatever was on it, went with her.

We build data and AI systems at Shipshape Data, and this is the conversation that opens most engagements. A client describes itself as somewhere between stage two and stage three on whichever maturity model they picked up at a conference, and when we look at what is actually running against real customers with nobody watching the dashboard for typos, a fair amount of it is stage one wearing stage three's slide deck. Nobody is lying, exactly. They are counting the wrong thing.

This guide sets out what the stages of AI maturity actually look like on the ground rather than in the diagram, how to score your own organisation without flattering it, and why the gap between what most companies claim and what they have built almost always comes down to one thing: mistaking a pilot that worked once for a capability that works every time.

The stages, and what each one actually looks like

Maturity models are everywhere now, five boxes and an arrow, usually sold alongside a consulting engagement that happens to start at whichever box the buyer is closest to. Some frameworks split this into five stages, some into four, one client of ours was handed a seven-stage version by a vendor who, unsurprisingly, positioned their own tool as the thing that carries you through stages three to five. The specific model matters less than getting honest about which stage you are actually in, and that takes more than reading the box labels. Here is what each stage looks like once you are inside it rather than presenting it at a conference.

Stage one: ad hoc experimentation

Someone technical, usually one or two people, is trying things. A ChatGPT wrapper here, a scikit-learn model there, built on whatever data they could get their hands on without asking permission through the proper channel. There is no shared infrastructure and no shared standard for what "good" looks like. Success is measured by whether the demo worked in the room. This stage is not a failure state. Nearly every organisation starts here and the fastest route through it is letting a handful of curious people experiment freely, before anyone tries to govern something that does not yet exist.

The tell for stage one is not the lack of polish, it is the lack of a paper trail. Nobody can say how many things have been tried, what happened to the ones that were quietly dropped, or where the data actually came from. That is fine at this size. It stops being fine the moment two people in different departments start building the same thing without knowing it.

Stage two: repeatable pilots

The organisation has done this enough times that a pattern has formed. There is a rough process for picking use cases, a preferred cloud platform, maybe a data science team of three or four rather than one enthusiast. Pilots get built with an actual sponsor and an actual success metric, decided in advance rather than invented afterwards to match whatever the model produced. Most pilots at this stage still die quietly. That is normal and, honestly, healthy. What is not healthy is when nobody tracks which ones died or why, so the same idea gets pitched again eighteen months later by someone who was not in the room the first time.

Stage three: managed production

This is the stage most companies claim and the fewest actually occupy. A model in production means it runs without a person babysitting it, it has an owner whose job description includes keeping it running, someone gets paged when it degrades, and there is a defined path for retraining or retiring it. The data pipeline feeding it is monitored, not assumed to be fine. The organisation can answer what happens on a Tuesday morning when the model is wrong and a customer is affected, and that answer is a process rather than a shrug.

Stage four: scaled and governed

Multiple production models exist across more than one business function, and there is shared infrastructure underneath them rather than each team building its own from scratch. A governance function exists that can say no to a use case on risk grounds and make it stick. Model performance gets reviewed on a cadence, not only when someone complains. Very few organisations we meet are genuinely here, and the ones that claim to be are worth a harder look than the ones that admit they are not. Genuine stage four is quiet. Nobody is presenting about it, because it has stopped being a project and become how the place runs.

Why the pilot count is the wrong thing to measure

Ask most leadership teams how their AI programme is going and the answer arrives as a number of pilots. Twelve this year, twenty last year, forty since the programme started. It is an easy number to report upward because it always goes up. It also tells you almost nothing about capability, because a pilot and a capability are answering completely different questions.

A pilot answers: can this work, once, for this data set, with this team watching it closely and quietly fixing whatever breaks. A capability answers: does this keep working when the team moves on to the next thing, when the input data drifts, when nobody remembers exactly how the demo was set up. Those are not points on the same scale. A hundred successful pilots that never survive contact with someone else's Tuesday is not a maturity ladder, it is the same rung repeated a hundred times.

The demo that never has to survive the real world

Pilots get built on curated data, run for an audience who wants them to succeed, and evaluated against whatever the person building them decided counted as good. None of that is dishonest. It is just a different, easier environment than production, and the skills that make a pilot land in a demo are not the same skills that keep a model correct six months later against data nobody cleaned first.

A pilot proves a model can work once, in a room, on data someone tidied up for the occasion. Production asks whether it still works after that person has gone on holiday.

The honest test of maturity is not how many pilots exist. It is how many of them are still running, unattended, a year later, and whether anyone would notice if one quietly stopped working.

Counting pilots rewards the wrong behaviour

Once a leadership team starts reporting pilot volume as a proxy for progress, teams optimise for pilot volume. That is a perfectly rational response to how they are being measured, and it produces exactly the pattern we see most often: a wide, shallow pile of proofs of concept, each one abandoned the moment it stops being novel, none of them ever asked to survive contact with a real customer on a bad day. A programme with three deployed, monitored, unglamorous models is further along than one with thirty pilots and nothing running.

Why most self-assessments come back a stage too high

Hand a leadership team a maturity questionnaire and watch what happens. Almost everyone scores their organisation a stage higher than an outside audit would. This is not really vanity, though vanity plays a part. It is mostly a structural problem with how the question gets answered.

Whoever is in the room decides the score

The people filling in a maturity self-assessment are usually the ones who championed the AI programme, and they are scoring the best examples they can think of rather than the median. One genuinely well-run production model gets used to answer a question about the whole organisation, while the eleven pilots stuck in someone's notebook do not come up, because nobody in the room is responsible for them and nobody wants to be the one bringing bad news to a workshop.

Ambition gets counted as capability

There is also a simpler mechanism at work. People answer maturity questions with where they are heading rather than where they have landed. A team that has agreed, in principle, to build a model governance function will often score itself as though that function already exists, because the intention feels close enough to the fact. It is not close enough. A model governance function that exists as a slide is worth exactly nothing the first time a regulator, or a customer, asks a hard question about a decision the model made.

We have sat in workshops where a room scored itself a clean stage three on paper, and the same room, twenty minutes later, could not agree on who owned the one model they pointed to as evidence. That gap between the score and the follow-up question is usually where the real answer lives. It is worth asking the follow-up question before trusting the score.

Best for running the self-assessment as an audit of what is actually deployed and unattended today, not a survey of intentions: pull up production logs and uptime records before anyone opens the questionnaire. Watch for: letting the same person who built the flagship pilot also score the organisation. Ask someone from finance or operations instead, someone with no reason to round up.

What production actually demands that a pilot never had to answer

The jump from stage two to stage three is where most AI programmes stall, and it stalls for reasons that have very little to do with the model itself. A pilot that scores well on accuracy can still be nowhere near ready for production, because production is asking a different set of questions entirely.

Someone has to own it after the excitement fades

Every pilot has a champion during the build. Few pilots have an owner for the eighteen months after launch, and that gap is where most of them quietly die or, worse, keep running wrong without anyone noticing. Production maturity means naming that owner before the pilot is even approved, not after it has already started drifting.

Monitoring has to exist before anything breaks, not after

A model that works today can be wrong in three months without a single line of code changing, because the world it was trained on has moved and nobody told the model. Catching that drift before a customer does is one of the clearest, most checkable markers of whether an organisation has actually reached stage three, and it is also where the gap between a pilot and a proper model deployment shows up first. A pilot gets deployed once, looked at, and left. A production model gets deployed, watched, and quietly redeployed as the data underneath it moves.

The infrastructure has to be built once, not per project

At stage two, every pilot tends to reinvent its own plumbing: its own data extraction, its own deployment path, its own way of tracking what changed and when. That is fine for one pilot. It becomes expensive fast once there are six of them, each with a slightly different way of doing the same job. Proper MLOps practice is what turns six bespoke builds into one repeatable pipeline, and its absence is usually the real reason a "scaling" programme is actually just running the same manual process in parallel a dozen times.

How to assess where you honestly stand

Skip the questionnaire, at least as the starting point, and start with a list. Every AI initiative the organisation has ever funded, one line each: what it was meant to do, whether it is still running today, who owns it now, and when someone last checked whether it still works. This exercise alone tends to be more revealing than any maturity framework, because most organisations discover they cannot actually finish the list. Nobody kept it.

Count what is unattended, not what exists

The number that matters is not how many models have ever been built. It is how many are running right now with no human checking the output before it reaches a customer or a decision, and have been for at least six months without incident. That is a small, blunt, honest number, and it is usually far smaller than the pilot count on the board slide.

Ask who gets paged

A genuinely useful question, asked plainly: if this model started producing wrong answers at two in the morning, whose phone would ring? If the honest answer is nobody's, the model is not in production in any meaningful sense, whatever the deployment dashboard says. This is the same discipline that proper AI governance is meant to formalise, and its absence is usually the clearest sign that maturity has been assumed rather than built.

The trap of chasing the next stage too soon

Once an organisation admits it is behind where it claimed, the instinct is to leap. Announce a stage four governance function, hire a head of AI, buy a platform that promises to scale everything at once. We have watched this go wrong more times than we would like, because governance built on top of two shaky pilots is theatre, not capability, and a platform bought to scale AI cannot fix the fact that the underlying data was never mastered in the first place.

Maturity, in practice, is closer to fitness than to a certificate. You cannot skip from occasional jogging to marathon shape by buying better shoes. Each stage has to actually be lived in, not announced, before the next one is worth attempting. A stage two organisation that spends a year making its handful of pilots genuinely unattended and reliable is in a stronger position than a stage two organisation that spends the same year writing a stage four governance charter nobody enforces. The unglamorous work is usually a data strategy that actually gets followed, plus the patience to let one thing work properly before starting the next.

There is a version of this we see often enough to name it: the governance function hired to sit above a programme that has nothing underneath it yet. A head of responsible AI arrives, writes a policy document, sets up a review board, and then discovers there is exactly one model to review, and it is a pilot that has not run since the person who built it changed teams. The policy was not wasted work, but it was built in the wrong order. Governance earns its keep once there is something real to govern. Before that, it is paperwork looking for a job.

Where to start

Build the list first: every model or pilot ever funded, its current state, its owner, the date anyone last checked it. That single document will tell you more about your real stage than any framework, and it usually takes an uncomfortable afternoon rather than a consulting engagement to produce.

Then pick the pilot closest to being genuinely unattended, the one where the gap between "it works in the demo" and "it runs on its own without anyone watching" is smallest, and close that gap completely before starting anything new. Name an owner. Put monitoring in front of it. Decide, on paper, who gets paged if it goes wrong. That is one production capability, honestly earned, and it is worth more than another twelve pilots that never leave the notebook.

If you want an outside, unflattering read on where your organisation actually sits, and what it would take to close the gap to real production capability, talk to us.

Start at your core.

Tell us where your data is today and what you want AI to do. We will come back with a straight answer on what your foundation needs and where the quickest real win is.

Talk to us