Data governance & compliance

12 data governance best practices to implement in 2026

A model got trained on customer records nobody could quite trace back to source. When someone in the room asked why it had flagged a loyal customer as high risk, nobody had an answer they trusted. That is not a modelling failure. That is a governance failure, and it is the one we run into most often when an AI pilot stalls before it reaches production.

We build data and AI systems at Shipshape Data, and the pattern repeats across nearly every client we work with. The algorithms are rarely the problem. The data underneath them is unreliable, poorly owned, or impossible to find, and no amount of model tuning fixes that. Companies that get real returns from AI have almost always made data governance a genuine priority rather than a policy document nobody reads.

None of this needs to be bureaucratic. Good governance removes friction rather than adding it. It is the reason a data scientist can trust a table without emailing three people first, and the reason a compliance officer can answer a regulator's question in an afternoon instead of a fortnight. Below are twelve practices worth putting in place this year, in roughly the order we would tackle them with a client starting from nothing. Some you can start this week. A couple take months. Take what applies and skip the rest, because not every organisation needs all twelve at once.

1. Start with an AI readiness assessment

Before you write a single policy or appoint a steward, find out where you actually stand. An AI readiness assessment gives you a factual baseline: what your data estate looks like, where governance is already breaking down, and what capability you have in-house to fix it. Skip this step, as most organisations do, and you end up writing policies for problems you have not actually confirmed exist.

We run this as a focused, time-boxed piece of work rather than an open-ended audit, typically about a month. That structure matters. Long assessments drift, produce hundred-page decks nobody reads, and lose the momentum that got them funded in the first place.

  • Which datasets actually drive revenue or compliance exposure, and which just take up space
  • Where duplicated or conflicting data lives across systems that were never meant to talk to each other
  • Whether your infrastructure can support real model training and deployment, not just a demo
  • Which teams already govern data informally, and which operate with no controls at all

Quantify what you find wherever you can. Hours spent reconciling two versions of the same customer table, revenue lost to a bad segmentation, the fine you would face if a regulator asked the wrong question this afternoon. A finding backed by a number gets funded. A finding that just says "data quality is poor" gets filed.

Best for any organisation about to invest seriously in AI without a clear picture of what it is building on. Watch for: assessments run as sales pitches dressed up as diagnostics; ask what evidence they actually collect, not just what slides they produce. We offer a free AI Readiness Assessment for exactly this reason.

2. Tie governance to business outcomes, not policy for its own sake

Governance that floats free of any business outcome gets ignored, and it deserves to be. Anchor it instead to something specific: launching an AI product, hitting a GDPR deadline, cutting customer churn. Once you can name the outcome, you can work backwards to the datasets, owners and standards that outcome actually depends on.

Vague goals do not force any useful decisions. "Improve data quality" tells nobody what to do on a Tuesday morning. "Cut churn by 15 percent using a predictive model by Q3" tells you exactly which dataset needs an owner, which fields need cleaning first, and which compliance checks apply.

  • Data must be fit for the AI use case it is going into, not just technically present
  • Privacy outranks convenience, every time, with no exceptions negotiated at 5pm on a Friday
  • Lineage must be traceable for anything that touches a regulated decision

Principles like these do a quiet job: they let a team resolve a disagreement on its own instead of escalating every judgement call to a committee. Link governance explicitly to the AI lifecycle too, collection, training, deployment, monitoring, so training data quality and output validation sit inside the same framework as everything else, rather than off in an AI ethics document nobody in engineering has read.

3. Get a named executive behind it, with real decision rights

Governance without a senior sponsor is theatre. You need someone with the authority to resolve a fight between two departments over who owns a dataset, and the credibility to make policy stick when a team decides the deadline matters more than the process. Without that person, expect shadow systems: spreadsheets nobody logged, exports nobody approved, the usual workarounds that appear whenever the official route is slower than the unofficial one.

Write decision rights down properly. "The data team decides" resolves nothing when there is a dispute at 4pm. Better: the Chief Data Officer signs off on external data sharing, domain owners control access to their own datasets, and everyone knows which of the two applies before the disagreement happens.

Keep the council small

A governance council of six to eight senior people, meeting monthly, achieves more than a committee of twenty meeting weekly. Give it real decisions to make, major policy changes, cross-domain conflicts, budget, and keep it out of routine approvals that would bury it in minutes within a quarter. Report to it. Do not run every decision through it.

4. Define who owns, stewards, and looks after each dataset

Vague ownership is the single fastest way to kill governance momentum. When nobody is quite sure who is responsible for a dataset, everyone assumes somebody else has it covered, right up until an audit or an outage proves otherwise. Organise ownership around business domains rather than technology: the commercial director owns the definition of "customer," a steward on the CRM team handles the day-to-day quality checks.

  • What the role decides, not just what it "ensures" (vague verbs are how you end up with unaccountable roles)
  • How many hours a month it realistically takes, so people can say yes with open eyes
  • Who it escalates to when something is outside its authority
  • What tools and templates it gets, because asking someone to police quality with no dashboard is asking them to fail

Build a RACI matrix across your governance activities, approving access, certifying quality, managing retention, and check two things: that every activity has exactly one accountable person, and that no single person is carrying five roles because nobody else volunteered. Then actually onboard your stewards. A ten-minute intro email is not training.

5. Write a policy set people will actually read

A framework without written policies is a set of good intentions. Write down how data gets created, accessed, shared and deleted, and keep the document short enough that someone reads it in one sitting.

  • A classification policy defining sensitivity levels
  • An access control policy stating who can view or change what
  • A retention policy setting how long you keep records before deletion
  • A data quality policy with minimum standards for completeness and accuracy
  • An AI usage policy covering model development and deployment

Make every principle measurable. "Data must be accurate" means nothing to an engineer building an alert. "Customer records must match source systems within 24 hours, with no nulls in mandatory fields" means something. Keep each policy to two or three pages, written for the person who has to follow it rather than the lawyer who might one day review it, and build a genuine exception process: business justification, a mitigation plan, an expiry date, reviewed quarterly so exceptions do not quietly become the new policy.

6. Find out what data you actually have, including the unstructured mess

You cannot govern what you have not found. Most organisations have a reasonable picture of their core databases and a near-total blind spot everywhere else: the SaaS tool finance signed up for without telling IT, the shared drive full of contracts, the chat channel where someone pasted a spreadsheet of customer emails eighteen months ago. Unstructured content like this often carries the most sensitive information in the business and the least oversight.

Start with the systems IT already manages, then expand outward to cloud platforms and SaaS applications where business teams quietly build their own shadow estate. Automated discovery tools help here, scanning storage and APIs for datasets nobody remembered to log. For each one, record where it lives, roughly how big it is, and why it exists.

Classify what you find by sensitivity, not just by system: public, internal, confidential, restricted, weighed against both regulatory exposure and what it would cost you if it leaked or got deleted by accident. Extend the same classification thinking to unstructured data, documents, emails, chat exports, because retention rules should apply there too, not just in the warehouse.

7. Give everyone the same definitions

Marketing's "active customer" and finance's "active customer" rarely mean the same thing, and the gap between them wastes far more time than anyone admits. A shared glossary, consistent metadata, and traceable lineage fix this, and they fix it cheaply compared with the meetings lost arguing over whose number is right.

Build the glossary around terms the business already uses in dashboards and board decks, not the terms engineers prefer. Define each one in plain language with a worked example, get sign-off from the people who actually rely on it, and require every critical dataset to carry a minimum set of metadata: an owner, a purpose, a refresh frequency, a sensitivity tag, a quality expectation. Block registration until those fields are filled in. It sounds heavy-handed until you see how fast "we'll add that later" becomes "nobody ever did."

Metadata turns datasets from mysterious files into documented assets people can actually use with confidence.

Trace lineage from source through every transformation to wherever it finally lands, automatically where your pipeline tools support it, and through an approved workflow diagram where they do not. That trail is what lets you investigate a quality problem in an afternoon instead of a fortnight, and what you hand an auditor when they ask where a number came from.

8. Build quality checks into the pipeline, not into a spreadsheet after the fact

A quality policy that nobody has automated is a wish. Translate it into rules a machine can check: order values greater than zero, customer emails containing an @ symbol, product codes matching the master reference list. Set a threshold that triggers an alert, an error rate above 2 percent, a spike in null values in a mandatory field, and put that check where the data enters the pipeline, not three reports downstream where the damage is already done.

Route failures somewhere useful. Log rejected records with a reason attached, so a steward can trace the fault back to its actual source instead of patching the same symptom every week. Assign a named owner to each alert category, and review, every quarter, whether the rules still match how the business and its AI models have actually changed. A rule written for last year's product will happily wave through this year's bad data.

9. Make privacy and access controls the default, not an afterthought

Build privacy and security in from the start rather than bolting them on once something has gone wrong. Grant access only to people who need it for their actual role, set permissions to expire automatically when a project ends or someone changes teams, and require an approval workflow with a proper audit trail for anything sensitive. Block bulk exports by default, and make someone justify the exception.

  • Retention periods set per dataset against real regulatory requirements, with automated deletion, including the backups and copies nobody remembers exist
  • Encryption at rest and in transit everywhere, with column-level encryption on the fields that matter most
  • Access logs monitored for the pattern that looks like a compromised account rather than an unusual but legitimate one
  • A quarterly check against UK GDPR requirements for consent, subject access requests and breach notification, plus whatever your sector adds on top

A retention policy enforced by hand fails the first time the team gets busy, and teams are always busy. Automate the deletion, not just the rule.

10. Govern the whole lifecycle, not just the storage layer

Governance that stops at "who can access this table" misses most of the actual risk. Data has a life: it gets created, transformed, shared, archived, and eventually it should be deleted, and each of those stages needs its own controls. Map the stages for your critical datasets, note who approves each handoff, and you will usually find gaps nobody had noticed, a dataset that gets copied into three tools with no owner tracking any of the copies, say.

Before you change a schema or a refresh schedule on anything critical, tell the people downstream first. Run an impact assessment, give consumers a testing environment to check nothing breaks, and keep a version history so you can roll back when, inevitably, something does. For data crossing borders or reaching third parties, add encryption and contractual protection, and actually reread those vendor agreements once a year rather than assuming they still hold.

11. Extend governance to cover models, prompts, and what comes out the other end

Standard data governance was not built with AI in mind. A model that generates content, makes a decision, or learns from how people use it needs its own layer of control: what data trained it, what prompts and knowledge bases shape its answers, and what checks run on its output before a customer sees it. This is what AI governance actually covers, and most organisations only discover the gap once something has already gone wrong.

Require sign-off before anything AI-driven reaches production if it touches customers, employees, or a regulated process. The trigger is usually one of three things: automated decision-making, personal data in the loop, or output that goes public. Track which datasets trained and evaluated each model using the same classification and lineage rules as everything else, and keep evaluation data genuinely separate from training data so you are testing against reality rather than the model's own assumptions.

  • Document every source feeding a retrieval-augmented or knowledge-base system, with a stated refresh frequency
  • Require approval before adding a new source that might smuggle in sensitive or unreliable content
  • Watch for knowledge-base drift that changes what the system says without anyone deciding it should
  • Actively look for shadow AI, the tool a team quietly adopted outside any approved process, before it causes an incident rather than after

AI governance that ignores the quality of its training data just produces a model that repeats every governance failure already sitting in your estate, faster.

12. Measure what matters, and actually act on it

A governance programme nobody measures eventually loses its funding, and it deserves to. Track the things leaders care about: time saved resolving data disputes, AI projects that actually reached production, fines avoided. Track the operational signals too: how long an access request takes to clear, how fast a quality issue gets fixed, what proportion of datasets carry complete metadata. Report the costs alongside the wins, quarterly, so the return is visible rather than assumed.

Run a lightweight maturity check once a year rather than a sprawling audit. Score a handful of dimensions, quality, access, lifecycle, against actual evidence rather than good intentions, and take the two or three worst gaps rather than writing a fifty-point improvement plan nobody will execute. Review policy twice yearly against what the metrics actually show, adjust in small steps, and pilot any real change with one team before rolling it out everywhere.

What to do next

You do not need all twelve running by January. Start with the readiness assessment, fix the two or three gaps it turns up that are already costing you money, and let those early wins build the case for the rest. Most of the organisations we work with get further with three practices done properly than twelve done half-heartedly.

If you want a straight answer on where your own gaps are before you spend anything on tooling, that is exactly the conversation to have. Talk to us and start with a clear picture of your data estate instead of a vendor's roadmap.

Start at your core.

Tell us where your data is today and what you want AI to do. We will come back with a straight answer on what your foundation needs and where the quickest real win is.

Talk to us