Data governance & compliance

Data privacy under GDPR and CCPA: a practical guide

A vendor demo for a support chatbot goes well, right up until someone in the room asks where the training data came from. Eighteen months of live chat transcripts, it turns out: names, order numbers, the odd health condition mentioned in passing to explain a return, and no record of anyone agreeing that any of it could be used to train anything. The demo continues. Nobody answers the question that day, and it should have been answered before the export ran, not after.

Data privacy under GDPR and CCPA is not one law wearing two names. It is two regulatory ideas that both talk about "privacy" and expect similar-looking paperwork, which is exactly why teams mix them up. GDPR assumes you need a reason before you touch someone's data. CCPA assumes you can touch it, but the person gets a say in whether you keep doing so. Confuse the two and you build a consent banner for a regime that barely uses consent as a lever, or ship a California-style opt-out link and call it done for Europe.

We build data and AI systems for clients at Shipshape Data, and privacy is usually where a promising AI project first meets a hard stop. Rarely because anyone was careless. Usually because the data was collected for something else entirely, years before anyone thought about training a model on it, and nobody checked what was promised at the point of collection. This guide covers what GDPR and CCPA require, where the regimes differ, how lawful basis and consent work, what data subject rights obligate you to do, the discipline of minimisation and retention, and what changes once personal data feeds or trains a model. None of it is legal advice, and nothing replaces a lawyer who knows your specific processing. It is the practitioner's map of where the obligations bite.

Why privacy failures cost more than the fine

GDPR fines get the headlines because the numbers are large: up to twenty million euros or four percent of global annual turnover, whichever is higher. That figure works as a deterrent, and it is also the least interesting cost, since the organisations hit with the maximum are rare. The common cost is duller and harder to see coming.

The fine is the least of it

A data subject access request takes real hours to answer properly: finding every system holding a person's data, checking what you can legally withhold, and producing an answer inside the statutory window. Do that badly a few times and word gets round. A breach notification clock runs from the moment you become aware, not from the moment you finish investigating, so the worst week of your year is also the week you have the least information.

A data subject access request is not an inconvenience your legal team quietly absorbs. It is a stranger asking you to produce, inside a month, exactly what you hold on them and where you got it, and most organisations discover on the first real request that they are not sure.

Enforcement reaches further down than the headlines suggest

The cases that make the news are the large ones. UK and EU regulators have also fined small businesses over things as ordinary as a badly configured marketing list or a CCTV system nobody registered a purpose for. CCPA works differently: California's Attorney General and Privacy Protection Agency enforce it directly, and consumers get a narrow private right of action for certain breaches, with statutory damages set between roughly a hundred and seven hundred and fifty dollars per incident. Multiply that across a customer list.

GDPR and CCPA solve different problems

Treat them as the same law in different vocabulary and you will build the wrong programme. They start from opposite defaults.

Opt-in against opt-out, the fork everything hangs on

GDPR's starting position is that processing personal data is not allowed unless you can point to one of six lawful grounds for it. No lawful basis, no processing, and that holds whether or not the person minds. CCPA starts from the other direction: a business can generally collect and use personal information as part of running the business, and the law's main lever is giving the consumer visibility and a way to say no, particularly to having data sold or shared. GDPR asks permission first. CCPA tells you what happened and lets you object. Confusing the two is how a company ends up asking for consent it never needed, or skipping the opt-out link that was the actual requirement.

Who has to comply, and who is watching

GDPR applies to anyone processing the personal data of people in the EU or UK, regardless of company size or location. A three-person startup with European customers is in scope on day one. CCPA only bites once a business crosses a threshold: broadly, revenue above a set figure, or handling a large number of California consumers' data, or deriving most revenue from selling it. A small business can sit outside CCPA entirely while being fully in scope for GDPR the moment one EU customer signs up.

Best for a single privacy programme meant to cover both regimes at once: design consent and disclosure flows to GDPR's stricter opt-in standard, since it clears CCPA's opt-out floor almost everywhere except the specific mechanics California requires regardless. Watch for: assuming a GDPR consent banner automatically satisfies CCPA, or the reverse. A "Do Not Sell or Share My Personal Information" link and a GDPR consent banner answer different legal questions, and a regulator on either side will notice if only one has actually been built properly.

Lawful basis comes before consent, not instead of it

Consent is the lawful basis everyone reaches for first, and it is often the wrong one. GDPR gives you six: consent, contract, legal obligation, vital interests, public task, and legitimate interests. Contract covers processing you need to deliver what someone signed up for. Legitimate interests covers processing that is proportionate and expected, provided you have run the balancing test and can show your interest does not override the person's. Consent is only needed when none of the others fit, and reaching for it as a default habit builds a fragile programme: it can be withdrawn at any moment, and whatever you built on top of it then has to stop.

What actually makes consent valid

Valid consent under GDPR has to be freely given, specific to a purpose, informed, and given through a clear affirmative action, not a pre-ticked box or silence, and as easy to withdraw as it was to give. Bundling consent to a privacy policy with consent to marketing tends not to hold up, because it is not really a free choice if refusing means you cannot use the service. Granular, per-purpose consent is more work to build and more honest about the ask.

Where CCPA draws the line differently

CCPA does not build its structure around lawful basis at all. Its default is notice and choice: tell people what you collect and why, and let them opt out of the sale or sharing of it. Opt-in consent only becomes mandatory in narrower cases, notably before selling a known minor's data, or certain sensitive categories such as precise geolocation, health data, or biometric identifiers. Treating it as equivalent to GDPR consent means overbuilding for California or underbuilding for Europe.

The rights people actually have

Both laws give individuals rights over their own data, and the rights are not identical in shape. GDPR grants a fuller set: access, rectification, erasure, restriction of processing, data portability, the right to object, and rights around automated decision-making, alongside the underlying right to be informed. Our guide to the eight GDPR data subject rights goes through each one, including how to action a request rather than just acknowledge it. CCPA's list is shorter and more transactional: know what has been collected, delete it, correct it, opt out of sale or sharing, and limit use of sensitive personal information, backed by a right not to be discriminated against for exercising any of them.

The response clock is shorter than people think

GDPR gives you one month to respond to most requests, extendable by two for genuinely complex cases, and you have to tell the person you are extending it. CCPA gives businesses forty five days, similarly extendable once. Both windows assume you can locate every place a person's data lives, which is what catches organisations out: answering a request is a search problem across every system that might hold a fragment of someone's record, and if nobody has mapped that, the first real request becomes the mapping exercise. Erasure is not absolute either. You can usually keep what a legal obligation, an ongoing contract, or a defence against a claim requires, provided you can explain which exception applies. A live record deleted while an untouched backup holds it for another eleven months is a genuine operational problem, not a loophole to hide behind.

Minimisation and retention: collecting less, keeping it for less time

Two GDPR principles do more practical work than any of the rights: data has to be adequate, relevant and limited to what is necessary, and it must not be kept longer than the purpose requires. Neither bans anything outright. Both make you justify a decision you would otherwise never have to justify, which is why they get ignored until an audit forces the question.

The instinct to keep everything is the enemy here

Most organisations collect more than the stated purpose needs, because an extra field on a form is easy to add and nobody wants to say no to "just in case" data. The same instinct shows up around retention: deleting data feels risky, keeping it feels safe, so records outlive their purpose by years. This is backwards. Unused personal data is unmanaged exposure with no offsetting benefit, and every extra field and extra year is more to defend if a regulator ever asks why you still have it. A retention policy nobody consults changes nothing. Shrinking real exposure means tying a retention period to each data category at the design stage and building deletion into the system rather than a manual task, which is uncomfortable, because it means someone has to admit a dataset kept for six years has not been queried in four.

What changes once the data trains or runs a model

Everything above still applies once AI enters the picture. It just gets harder to see.

Training data has to earn its own purpose

Data collected for one reason, running a support desk, processing an order, does not automatically carry a lawful basis for a second, different purpose: training a model. That reuse is new processing, and it needs its own justification, not a borrowed one from whatever consent or contract covered the original collection. This is not theoretical. Italy's data protection authority temporarily ordered ChatGPT blocked in March 2023 over exactly this kind of question, including whether the training data had a valid legal basis and whether users had been properly informed. The service returned once changes were made, and the episode previews the question any regulator eventually asks: where did this come from, and did anyone agree to this specific use.

The model remembers what the record forgot

Deleting a row from a database is straightforward. Deleting what a model learned from that row during training is not, because the information is no longer a discrete fact, it is distributed across weights that also encode millions of other records. Genuine machine unlearning, removing one person's influence without retraining the whole model, remains a hard technical problem, so the honest answer to "can you erase me from the model" is often "not without retraining it." Pseudonymising data before training helps, though large language models can memorise and occasionally reproduce snippets of their training data, so pseudonymisation alone is not always enough of a shield.

Sending personal data to someone else's model is still a transfer

Using a hosted large language model on customer data means that data is leaving your organisation for a processor, sometimes a sub-processor several layers deep, and if that vendor sits outside the UK or EU it is also a cross-border transfer needing its own legal mechanism, typically standard contractual clauses. None of that disappears because the interface is a chat window, not a database connection. Where the personal data is not actually needed, synthetic or de-identified alternatives are worth considering instead of defaulting to the real thing because it was already sitting there; our piece on synthetic data covers that trade-off.

Best for teams deciding whether a hosted LLM API can touch real customer data at all: check the vendor's data processing agreement, sub-processor list and retention terms before the pilot starts, not after it already works. Watch for: "we do not use your data to train our models" is a promise about that vendor's own training runs. It says nothing about logging, sub-processors, or how long your prompts sit in a support queue somewhere, and that gap is usually where the actual exposure lives.

Automated decisions still need a human somewhere

GDPR gives people a right not to be subject to a decision based solely on automated processing, including profiling, where the decision has legal or similarly significant effects, such as a loan refusal or a hiring rejection. A model scoring an application, with a human rubber-stamping the score without really looking at it, does not satisfy this. California is moving the same way: its automated decision-making rules under the CPRA build a comparable notice-and-opt-out requirement. Wherever a model's output changes what happens to a real person, somebody accountable needs to be able to explain, and where the law requires it, override that output.

The unglamorous mechanics that keep you compliant

None of the above stays true on its own. It needs a small set of unglamorous, recurring admin behind it, and skipping this is how a programme that looks fine on paper turns out to have nobody who can answer a regulator's first question.

Assess the risk before you build, not after

GDPR requires a data protection impact assessment before starting processing likely to result in high risk to individuals, which in practice covers most serious AI projects touching personal data: large-scale profiling, systematic monitoring, or sensitive categories at real scale. Our guide to DPIAs covers what one needs to contain. Doing this once the model is in production turns a design conversation into a retrofit, and retrofits are where teams discover the architecture cannot support the safeguard the assessment needs.

Records, contracts, and the seventy-two-hour clock

GDPR expects a record of processing activities, kept current rather than written once for an audit and forgotten. Every processor handling personal data on your behalf, including any AI vendor, needs a data processing agreement setting out what it can and cannot do with that data. If a breach happens, the clock starts at seventy-two hours to notify the relevant authority, counted from when you became aware, not when the investigation concludes. CCPA imposes no such fixed window generally, though California's own breach law still applies.

Where to start

Start by finding out what you actually have, not what the last data map claims you have. Most organisations that think they know where personal data lives are wrong about at least one system, usually a spreadsheet, a support tool, or a training export somebody ran eighteen months ago and forgot to log. That inventory is dull work, and it is the only foundation the rest of this rests on. Then work out, dataset by dataset, what lawful basis or CCPA disclosure applies, whether the retention period matches a real purpose rather than habit, and whether anything feeding an AI system was collected for that specific use. Fix the gaps before the next model gets built on top of them, not after.

Privacy compliance is not a document you finish once. It is a discipline you keep applying as the data and the systems around it change, and that is doubly true once AI is involved, because the questions get harder to answer and the mistakes get harder to see. If you are building an AI system on data you are not fully sure is clean on the privacy side, or you want a straight assessment of where GDPR or CCPA bite for your specific processing, talk to us. We would rather help you find the gap before a data subject access request finds it for you.

Start at your core.

Tell us where your data is today and what you want AI to do. We will come back with a straight answer on what your foundation needs and where the quickest real win is.

Talk to us