Plans
Learn Library

AI Grew Up Eating Data — But That Meal Didn't Come Free

An educational piece explaining how GDPR, CCPA, and PIPL affect AI products, covering cross-border data rules, deletion and bias challenges, plus privacy by design and privacy-preserving techniques such as federated learning, differential privacy, and secure multi-party computation.

ai-marketing
2026-09-18SupaMarketers8 min read

A few days ago, a friend of mine who builds AI apps invited me out for tea.

He had just hooked his product up to a large AI model and fed it the user data he'd been stockpiling for years, and his recommendation accuracy was climbing fast. Mid-story, caught up in the excitement, he slapped the table: "We're set now."

I asked him: that data belongs to your users, right? When you took it, did you ask them?

He froze, teacup in hand.

I told him: you're enjoying this meal. But it isn't a free one. Someone's keeping the tab.

Who's keeping the tab? Three laws. One in the European Union, one in the US state of California, one in China.

Let's take them one at a time.

First, See How Thick the Ledger Is

So what exactly is the GDPR?

The EU's General Data Protection Regulation, in force since 2018. Among privacy laws worldwide, it's the one most others have been copied from.

It sets a few hard rules: if you collect data, get consent, plainly and clearly; collect only as much as you'll use — no grabbing extra just because it's convenient; and users can inspect, correct, or delete their data at any time.

Delete is where it gets really fierce. Users can demand that you wipe their data entirely — what the industry calls the "right to be forgotten."

And if you break the rules? Fines. Whichever of two numbers is higher: €20 million, or 4% of your global annual revenue.

Note the word global. If you make 10 billion RMB a year, 4% of that is €40 million — over 300 million RMB.

California's CCPA (California Consumer Privacy Act) runs on a different logic. Its core concern: if your data is being sold, you must be told — and given a place to say "don't sell mine." That's why Californians keep seeing a link on websites: "Do Not Sell My Personal Information." Each violation costs $7,500.

China's PIPL (Personal Information Protection Law), effective November 2021, is the first law to give personal information protection the full treatment. Access, correction, deletion — it has them all. But it adds one more layer: data sovereignty. Data collected inside China must, in principle, stay inside China; to send it out, you have to meet specific conditions.

The fines are no joke either: up to 5% of annual revenue, plus possible revocation of your business license.

For a company making 1 billion RMB a year, the top-end fine is 50 million. That cut goes straight into the flesh.

One Set of Data, Three Rulebooks

Some readers may be asking: my company is governed by all three laws at once — whose rules do I follow?

That's exactly where the trouble lies. The three of them don't recognize one another.

Consider the thorniest issue of all: cross-border data flows.

In 2020, the EU's highest court stepped in and struck down the Privacy Shield — the data passport that had been passed around for years — ruling it invalid. The reason: US government surveillance could not adequately protect the data of EU users.

One ruling, and countless companies shuttling data back and forth across the Atlantic lost their legal channel overnight. All they could do was fall back on Standard Contractual Clauses (pre-approved contract templates that let companies move data legally), signing them one by one. Even Google got dragged into litigation for relying on those clauses to transfer data to the US.

Later, in 2023, a new framework arrived and the passport was patched. But after all that back-and-forth, everyone had learned their lesson: don't put all your eggs in one basket.

Add on top of that China and India, which both require sensitive data to be stored locally — servers built in-country, teams hired locally. The money just drains away.

So the standard playbook for multinationals today sounds almost sad: slice the data up by jurisdiction and store each pile separately.

The same data, three rulebooks, three times the cost.

When AI Steps In, the Awkwardness Doubles

Everything above is still just the problem for ordinary internet companies. Once AI steps in, the awkwardness doubles.

The first snag: AI's appetite and the "data minimization" principle are natural enemies.

For AI to perform well, it has to eat a lot. But GDPR says data collection must be minimized and purposes limited up front. Training a model to read medical images might take hundreds of thousands of patient records — and patient records are sensitive data. Without a patient's explicit yes, you can't so much as touch them.

The second snag: the "right to be forgotten" collides with the model.

A user says: delete my data. Fine — delete it from the database, a one-minute job.

But the data has already been trained into the model. A model is not a database; it has digested the data into hundreds of millions of parameters. The only way to spit it back out is to retrain from scratch.

One retraining run means millions in compute plus weeks of time. So when a deletion request comes in, do you take it? Take it, and the cost hurts; refuse, and you're breaking the law.

The third snag: the black box.

What's a black box? The model hands you an answer but can't tell you why.

In 2020, the scoring system of a major credit card company came under fire: with nearly identical profiles, women were getting visibly lower credit limits than men. Deliberate discrimination? Nobody could say for sure. Even the people who built the system could only see the inputs and the outputs.

The industry later came up with "explanation tools" like LIME and SHAP, which translate a model's judgments into something readable. Helpful. But still a long way from opening the black box completely.

And then there's Amazon's famous lesson: because the training data held more male résumés, the recruiting AI learned to score women down — and in the end the whole project was killed.

Bias that rides in with the data isn't digested by AI — it gets amplified and served up to you.

That's why the EU went on to pass the AI Act, already in force and phasing in: high-risk AI systems must have transparency, safety, and the rest squared away in advance.

So What Do You Do? Two Moves

Don't panic just yet. There is a way through.

Move one: draw privacy into the blueprint before work begins.

This idea has a name: privacy by design.

What does it mean? When you renovate an apartment, you run the wiring and plumbing first, then build the walls — you don't move in, trip the breakers every day, and then knock the walls down to start over.

Applied to an AI project: run a privacy risk assessment before you start, and trace the data flow from end to end; encrypt everywhere encryption can be used; and at every step, ask one more question — is this data truly something we must collect?

Move two: the data stays home — and still gets the job done.

These techniques go by the umbrella term privacy-preserving computation. Here are three.

First, federated learning.

The old way: each bank copies its customer data onto one server, and everyone trains a model together. All the data sitting out in the open.

Federated learning flips it around: the data stays home, and nobody moves it. What gets sent out is what the model has learned. A central server pools everyone's study notes, and the model is trained. The raw data never leaves the house.

FATE, open-sourced by WeBank, does exactly this. Banks and insurers use it to fight fraud and manage risk together — without ever opening their data to one another.

Second, differential privacy.

You add just the right pinch of noise to the statistical results.

How much? There's a strict mathematical standard: whether one person joins or leaves the dataset, you can't detect any difference. That way you can see what the crowd looks like, but you can never pinpoint what any one person looks like.

Apple made this famous: collecting iPhone users' habits — which emoji are most popular, which apps keep crashing. The statistics still get done, but no single piece of data can be traced back to a specific phone.

Third, secure multi-party computation.

Several institutions want to run the numbers together, but none of them wants to show their cards. The fix: each locks its data into an encrypted box, the computation is done on the boxes, and at the end only the result is opened.

When it's done, everyone knows the answer — and not a single card has been turned over.

Don't Assume This Is Someone Else's Problem

Finance has already made this playbook work end to end.

Fraud detection runs on AI watching in real time: a login from another city, wrong passwords typed again and again, a strange large transfer in the middle of the night — blocked on the spot. Identity verification leans on faces and voiceprints. Cross-border compliance gets routed around with federated learning and secure multi-party computation: the data never leaves the country — and the model still gets trained jointly.

Some will say: our company is small — nobody is watching our little pile of data.

One Facebook breach put the personal information of 533 million people up on the open internet.

533 million — more than the entire population of the United States. Unbelievable.

Meanwhile, another survey found that 87% of consumers consider data privacy a basic human right.

Users do care. They're just waiting for something to go wrong.

The ceiling on AI has never been compute. It's trust.

Back to my tea-drinking friend at the start. My advice to him came down to one line: while the project is still young, draw compliance into the blueprint. Don't wait for the fine to arrive before you go back and learn what GDPR even is.

May your AI eat its fill of data — and pass the audit.

Continue reading