AI governance & ethics

AI governance best practices: framework, policies and controls

A model goes live. Six months later a customer asks why they were turned down for a loan, and nobody on the team can fully reconstruct how the decision was made. That gap, between what your AI does and what you can prove about it, is what AI governance closes. Get it right and you can defend every deployment. Get it wrong and you find out during an audit, or worse, during a headline.

This article walks through what effective AI governance looks like: the frameworks worth adopting, the policies you need to write, and the controls that keep everything accountable once it's live. The aim is concrete steps, not vague principles.

Why AI governance matters now

The conversation has moved on. A few years ago, most boards treated AI governance as something to revisit once the technology settled down. That patience has run out. Regulators have caught up, AI failures have made the news often enough that risk committees now ask pointed questions, and organisations without a defined approach to building and monitoring AI carry exposure well beyond a technical glitch.

Regulation isn't optional any more

The EU AI Act phased in through 2024 and 2025, and its toughest provisions for high-risk systems are now fully in force. If you operate in or sell into the European market, you carry binding obligations around transparency, human oversight and documentation for any AI system touching employment, credit or other consequential decisions. The UK has hardened too. The FCA, the ICO and the CQC have moved from general principles toward sector-specific expectations, and each has shown willingness to act on them. Non-compliance is not just a fine risk. It can mean enforcement action, restrictions on what you're allowed to run, and public scrutiny that outlasts the incident itself.

Boards have started asking questions their technical teams need real answers to. How do we know this model performs fairly across different groups? Who signed off on this deployment? Good governance means those answers already exist before anyone has to ask for them.

AI now sits inside high-stakes decisions

This isn't only a large-enterprise problem. Mid-sized firms run AI on customer scoring, contract review and forecasting, often with little more structure than someone on the data team building it because it seemed to work. The moment a system moves from suggesting things to deciding things, the cost of getting it wrong changes shape. A biased model doesn't produce one bad output. It can discriminate quietly, at scale, across thousands of decisions before anyone spots the pattern.

Poor governance compounds inside the data itself, too. Models trained on inconsistent or incomplete data degrade over time, and without monitoring you won't catch that drift until it has already done damage. We say this often to clients: the model is rarely the weak point. The data underneath it usually is.

Retrofitting governance costs more than building it in

Most organisations discover they need governance only after something has gone wrong: a breach through a poorly secured pipeline, a model whose outputs nobody can explain to a client, an automated process that quietly breaches a contract. Adding controls after the fact is much harder than building them in from the start. You're not just writing documentation retroactively. You're often re-architecting pipelines and retraining models while the business keeps running on the thing you're trying to fix.

The organisations that get ahead of this build the unglamorous infrastructure early: named owners, risk classification, audit trails, before they scale AI across the business. That investment pays for itself the first time it prevents a serious incident.

What good governance actually looks like

Good AI governance isn't a policy document sitting in a shared drive. It's a working system of people, process and technical controls that keeps AI reliable, explainable and pointed at what your organisation wants it to do. Done well, it's almost invisible: things run smoothly, decisions leave a trail. Done badly, you feel it immediately, usually at the worst possible moment.

It's proportionate, not uniform

Not every AI system deserves the same scrutiny, and treating them all the same wastes effort in one direction while leaving real exposure open in the other. A tool that recommends product categories sits in a completely different risk environment to a model that influences credit decisions or flags someone for investigation. Applying the same checklist to both slows the low-risk work down without protecting the high-risk work.

Layer your controls by consequence instead. Low-risk tools need basic documentation, a named owner and a periodic review. High-risk systems need a formal impact assessment, a human oversight mechanism, a clear audit trail and sign-off before anything ships. Write these tiers down so teams know what's expected of them before they start building, not halfway through.

It assigns ownership clearly

Ambiguity about who owns an AI system is one of the most common failure points we see. Everyone assumes someone else is watching model performance. The documentation exists somewhere but nobody maintains it. Then something breaks and accountability turns into a game of pointing at other departments.

Name an owner for every system, from the first line of code through to the day it gets switched off, and make that ownership mean something real: monitoring, incident response, keeping the system inside its approved risk tier. Ownership shouldn't sit with one function alone. It spans data teams, product, legal and whichever business unit is actually using the thing.

It runs continuously

Governance that only shows up at deployment misses most of the risk. Models drift. Data quality erodes quietly. Rules change. Treating governance as a one-time gate rather than an ongoing job is how organisations end up compliant on paper and exposed in practice.

Running it properly means regular performance reviews, automated monitoring wherever that's feasible, and a clear process for escalating anything that falls outside expected bounds. It also means the programme changes as your AI portfolio grows: new controls when you move into higher-risk territory, and retiring the processes that no longer reflect what you're actually running.

How to choose a governance framework

No framework fits every organisation out of the box, and picking one without weighing your risk profile, regulatory context and operational maturity is a common mistake. A good framework gives you a structured starting point. It's a tool though, not a finished answer.

Map your risk profile first

Before you compare frameworks, work out what you're actually trying to govern. A firm running AI in financial services, where the FCA expects explainability and audit trails, starts from a different baseline than one using AI purely for internal knowledge search. Catalogue your current and planned AI systems, then classify each by what a failure would actually cost you: regulatory, financial, reputational, operational. That tells you how much rigour you genuinely need, and which frameworks are worth your time.

The main frameworks worth knowing

A handful of established frameworks give you a solid base. The NIST AI Risk Management Framework is one of the most usable. It organises governance into four functions: Govern, Map, Measure and Manage, a structure that applies well to organisations building a programme from nothing.

ISO/IEC 42001 is worth a look if you want a certifiable, internationally recognised management system standard. It sits comfortably alongside ISO 27001 or ISO 9001 if you already run those. And if you operate in or sell into the EU, the AI Act's own risk classification approach should shape how you tier systems, whichever primary framework you run day to day.

Adapt rather than copy wholesale

No framework maps cleanly onto your existing structures, team size or technology stack, and forcing one to fit usually produces governance nobody follows. Treat any framework as a reference, not a script. Work out which parts apply directly, which need adjusting, and which you can set aside because they don't reflect the risks you carry.

Best for teams starting from nothing: pick the NIST AI RMF as your backbone and add ISO/IEC 42001 later if a client or regulator needs certification. Watch for: a framework adopted wholesale, with no adaptation to your own stack and team, tends to sit in a folder rather than change how anyone actually works.

Most mature programmes draw from two or three frameworks and build something bespoke around them. That takes longer up front, but it produces governance people actually follow, rather than governance that exists only in documentation nobody reads.

How to define policies and controls by risk

Policies without a risk-based structure are one of the most common failures we run into. An organisation writes one AI policy, applies it everywhere, and finds it either strangles low-stakes work or leaves the systems that matter most under-governed. Start from risk classification and build outward, so controls match what actually happens if something goes wrong.

Build a risk tiering system

Decide how many risk tiers you need and what puts a system into each one. Three or four tiers works for most organisations, running from minimal at the base to critical, or prohibited, at the top. Base the criteria on how directly the system affects people, whether it touches a regulated decision, how many decisions it makes, and how reversible a bad outcome would be.

Build a simple tool, a questionnaire or a decision matrix, that lets product owners and data teams assign a tier during scoping, before real development starts. Do this early and governance requirements are known up front rather than discovered after the fact.

Set policies per tier

Once tiers exist, map specific requirements to each one rather than keeping a single blanket policy. Minimal-risk systems might need basic documentation, a named owner and an annual review. High-risk systems need a formal impact assessment, documented data provenance, explainability standards and a mandatory human review step before anything consequential gets acted on.

Be specific about what each tier requires. Teams shouldn't have to guess their obligations. Write policies with clear thresholds, including what triggers an escalation to a higher tier if a project's scope changes partway through.

Build controls into the workflow

Policies only work when the controls sit inside your normal development process rather than getting checked the week before launch, when nobody has time to fix what they find. At design sign-off, pre-deployment review, and post-launch monitoring, teams should confirm the controls for their tier are actually in place and documented. Tie these checkpoints to whatever project process you already run, so following the policy is the easy path rather than a track people quietly route around.

How to set roles, decision rights and a RACI

Undefined ownership is one of the fastest ways to undo an otherwise solid governance programme. When nobody is clearly accountable for a system's ongoing performance, monitoring lapses, incidents don't get escalated, and compliance gaps widen unnoticed. Role clarity isn't paperwork. It's what keeps governance running day to day rather than sitting dormant in a document.

Name an owner for every system

Every AI system needs a named owner, accountable from initial scoping through to the day it's switched off. That's not a symbolic title. The owner makes sure the system stays inside its approved risk tier, that monitoring is actually running and being reviewed, and that any change to scope or data inputs goes through proper sign-off.

Ownership usually sits with the business unit that benefits from the system, not the team that built it. Technical teams answer for build quality. The business owner answers for what the system actually does in production, and for what happens when it doesn't behave.

Set decision rights across functions

Once owners are named, be explicit about which decisions a function can make on its own and which need sign-off from elsewhere. Shipping a low-risk internal tool might sit entirely with a product team. Deploying a model that shapes customer-facing decisions in a regulated context needs input from legal, compliance and senior leadership first.

Ambiguous decision rights don't just slow things down. They create the conditions where the biggest decisions get made by default instead of on purpose.

Document these thresholds explicitly, so teams know in advance when to escalate rather than finding out after a deployment has already gone live.

Build a RACI people actually use

A RACI matrix assigns four roles to each governance activity: Responsible, Accountable, Consulted, Informed. The value isn't in covering every scenario. It's in producing something specific enough that people can act on without a meeting to interpret it.

For each stage of the AI lifecycle, name who does the work, who answers if it goes wrong, who needs consulting before a decision is made, and who just needs to know once it's done. Focus the matrix on the points where ambiguity actually bites: tier classification, deployment approval, incident escalation. Don't try to map every activity across every team.

How to govern the AI lifecycle end to end

Governance applied only at deployment misses most of the risk surface. Every stage of a system's life, from scoping through to retirement, carries its own risks and needs its own controls. Treating governance as something that runs the whole way through, rather than a single checkpoint, is one of the more practical shifts an organisation can make.

Govern the build before it starts

The decisions that shape how governable a system ends up being get made early: what data it trains on, what it will decide, what success looks like in measurable terms. Set governance requirements at scoping stage and teams know what documentation and approval steps apply before they write a line of code.

A short pre-development checklist works well here: the system's intended purpose, the data it needs, its risk tier, and the name of its business owner. That document becomes the baseline you measure later changes against, and it saves the usual scramble to back-fill documentation the week before launch.

Review governance at each transition

As a system moves from development into testing, staging and production, each transition should need a formal sign-off confirming the controls for its tier are actually there. This isn't bureaucracy for its own sake. It's catching gaps while they're still cheap to fix.

Define specific checks per stage, based on tier. A high-risk system moving into production should need documented explainability standards, confirmed human oversight and a rollback plan that's actually been tested, not just written down. Lower-risk systems need lighter checks, but verify readiness before you move on regardless.

Plan decommissioning from day one

Most governance programmes focus heavily on deployment and monitoring and give little thought to how a system gets retired. That creates its own problems: models stay in production long past their useful life, and nobody owns shutting the thing down.

When you first classify a system, write the decommissioning plan at the same time: who decides when it's retired, what happens to the data, what obligations survive after it stops running. Do this at the start and retirement becomes a managed process instead of something that happens by neglect, usually years after anyone remembers why the system was built.

How to monitor, audit and handle incidents

Deploying a system isn't the finish line for governance. It's where the ongoing work actually starts. Models degrade. Data distributions shift. Edge cases turn up that nobody saw during development. Without active monitoring, real audits and a tested incident response process, governance lives in your documentation and nowhere near your actual operations.

Set up continuous monitoring

Every system in production needs performance metrics tracked automatically, not reviewed once a quarter if someone remembers. At minimum, track output quality, input data consistency, and whatever fairness indicators matter for that system's risk tier. Automated alerts tied to threshold breaches give you warning before a problem becomes an incident.

Build your monitoring around the risk tier you assigned at scoping. High-risk systems need close to real-time monitoring and tight alert thresholds. Lower-risk tools can run on a lighter cadence, but they still need a defined review schedule and a named person responsible for looking at the output.

Run structured audits

Monitoring catches performance problems. Audits check something different: whether your governance controls are working as designed. A proper audit checks whether the system still matches its original risk classification, whether the documentation is current, and whether human oversight is actually happening, rather than being quietly skipped when things get busy.

Run audits at a frequency tied to risk: at least annually for lower-risk systems, more often for anything touching regulated or high-impact decisions. Give audit responsibility to someone independent of the team that owns the system day to day. Self-certification isn't an audit.

Define the incident response process

When something goes wrong, how well you handle it depends on whether you worked out the process before the incident happened. Document a clear escalation path: who gets told, who has the authority to suspend a system, and what you're obliged to tell affected users, regulators or clients.

Test the process before you need it. A tabletop exercise, walking the team through a realistic failure, exposes gaps in your RACI and your communication plan while the stakes are still low enough to fix cheaply. Incident response designed under real pressure is almost always thinner than it should be.

Getting started

Organisations that treat governance as an operational function, rather than something bolted on to satisfy an audit, are the ones that scale AI without the incidents that set a programme back by months. Risk tiering, policy design, lifecycle controls, incident response: every piece here connects to the same outcome. AI you can stand behind in production, under scrutiny, over time.

Start by auditing what you already have running. Find the AI systems already in production without a named owner, a risk classification or any documented monitoring. That gap tells you where to start, building layer by layer, with whatever carries the most risk first.

If you want help building a governance programme that actually fits how your organisation works, rather than a template pulled off a shelf, talk to us.

Start at your core.

Tell us where your data is today and what you want AI to do. We will come back with a straight answer on what your foundation needs and where the quickest real win is.

Talk to us