Skip to main content

Responsible AI starts with responsible data

The hardest data decisions are not made during deployment. They are made during collection – and they are the ones you cannot take back.

AI model defects are usually fixable. You can retrain, fine-tune, add a rule, roll back to a previous version and do so much more. Data defects are frequently far less forgiving. A misstep, like assembling a dataset without a clear right basis or holding personal information that should never have been captured, becomes an inheritance. Every model trained on such data carries the flaw forward, and the only full remedy may be to start over from scratch.

In the United States, regulators have said this explicitly. The Federal Trade Commission’s remedy for Cambridge Analytica in 2019 and Rite Aid in 2023 was the destruction not only of the improperly collected data but of “any data, models, or algorithms derived in whole or in part therefrom. A photo-app company lost a working face-recognition model in the same way. Europe has no direct equivalent, although supervisory authorities can order erasure or suspend processing.

The scale of the systems in question may change the outcome, but not the principle. The largest settlements have led to the destruction of source libraries while leaving the trained models standing. Most organizations, however, don’t have the leverage for such favorable results.

Reframing the challenge

Data preparation is still treated as the unglamorous work preceding the real project. In practice, however, the decisions with the longest shelf-life are made during data collection: what was gathered, from where, under what terms, and what was recorded. A missing license is not a documentation gap that an audit can close; it is an absence of a fact.

Four shifts are changing the way decisionmakers should understand and approach responsible data collection where AI is involved, and a fifth shift is about to make the other four enforceable.

Not every data flaw can be fixed later.
Some flaws such as duplicates, malformed records and missing values. These can be caught by good governance practices. Mislabeled data can be costly but recoverable through relabelling and retraining. What cannot be recovered is what was never collected. If a dataset under-represents a population, a condition or an operating environment, it’s hard to invent the missing evidence.

An important distinction is between random and systematic error. Models tolerate randomness and degrade gracefully, but systematic errors are learned and adopted, because a consistently wrong artefact is indistinguishable from a real signal. Geirhos and colleagues described this shortcut learning in Nature Machine Intelligence: decision rules that perform well on standard benchmarks but fail to transfer to more challenging testing conditions.

Intellectual property: The bell cannot be unrung.
Teams under delivery pressure scrape data. It is fast, readily available and appears to be free – and it is the data decision most likely to be irreversible.

Two 2025 outcomes frame the exposure. A US settlement saw an AI developer agree to pay authors approximately USD 1.5 billion over works taken from pirated sources. In the UK, Getty Images v Stability AI turned on jurisdiction and storage: Getty withdrew its primary training claims because training occurred abroad, and the court rejected secondary infringement because the model did not store the works. Neither outcome found that scraping is safe.

Nor is the European position settled. The text and data mining exception permits mining unless rights are reserved in machine-readable form, and its limits are being drawn case by case. Hamburg’s appeal court held in Kneschke v LAION that a reservation written in ordinary language is not machine-readable, because no automated process can act on it. In GEMA v OpenAI, Munich held that the exception does not stretch to include content that a model has memorized and reproduced. Neither judgment is final, and the Court of Justice just heard its first generative-AI copyright case in March 2026.

The lawful basis most European teams rely on is still being defined.

Privacy must be designed into collection.
Anonymization is often scheduled as a downstream step. That sequencing is a flaw. The Court of Justice confirmed in EDPS v SRB (September 2025) that personal data is a relative concept but the organization holding the original always retains the means of re-identification. Therefore, anonymization must happen at data capture, which means identifying information is never stored.

Technique matters, and it’s often a question of privacy versus accuracy and value. A study of face obfuscation in ImageNet found that blurring costs about one percentage point of classification accuracy. That does not generalize later work on detection, and segmentation found that blurring and masking substantially degrade performance when whole bodies are obscured. To make matters worse, redaction is exactly the kind of consistent artefact a model will learn to deal with. Images with blacked-out regions weaken the model while leaving individuals recognizable from context.

Generating synthetic faces to replace real ones can help avoid both problems. The scene keeps its realism and its training value, while a degree of privacy is maintained. So, there is value in understanding the difference between subtracting information and substituting it, but neither is as effective as not collecting it to begin with.

Data is not just an asset; it’s an attack surface.
Security teams are well-versed at protecting data as an asset. However, where AI is concerned, data is a critical source of input, and inputs can be manipulated.

Data poisoning is the clearest example. Researchers showed that it’s possible to spend a low double-digit dollar figure and control a small fraction of widely used public image-text datasets. In the case they describe, even manipulating less than 1% of the data was enough to affect targeted behavior.

Extraction is kind of a mirror image of poisoning, but the risk is just as serious. Production language models have been shown to emit verbatim training data under adversarial prompting, so a record collected carelessly can resurface as a disclosure or breach much later – and there is no un-training it.

To make things more difficult, a model’s corpus is also no longer frozen at implementation. Retrieval-augmentation and agentic systems pull documents in at query time, so new batches of untrusted content can reach models without touching the training set, and retrieval indexes routinely flatten the access controls of their sources. NIST’s adversarial machine learning taxonomy and the OWASP Top 10 for LLM Applications treat both issues as first-order risks.

The control that matters is lineage. When a dataset turns out to be poisoned, wrongly licensed or over-collected, the first thing to do is identify which models these are and which indexes they have reached. Without this information, there’s no clear way to contain the damage, much less mitigate the root problem, yet few organizations have the documentation they need to answer the question in a timely manner.

Baseline expectations are rising, and they call for evidence.
Until recently, good data governance was an assertion. It’s now becoming a specification.
Article 10 of the EU AI Act requires training, validation, and testing data for high-risk systems to sit under documented governance, covering factors such as origin and original collection purpose, preparation, assumptions, suitability, bias examination and mitigation and gaps. The Digital Omnibus, in force since July 2026, moved those obligations to December 2027, but that’s only a brief reprieve. The deadline moved; the difficulty did not.

Two European drafts will determine how demanding the requirements have become. prEN 18284, on the quality and governance of datasets, is written against Article 10 and spans data acquisition through retention; it also refers normatively to prEN 18283 on bias management, so bias handling becomes a dependency of dataset quality rather than a distinct exercise.

AI datasets will require a specified file containing the source and license for each batch, collection date and consent basis, the labelling instructions and who applied them, measured class balance across the groups the system will meet in service, and a retention schedule. ISO/IEC 5259-2 already supplies the vocabulary for the measurable parts, including identifiability – an explicit measure of residual personal data risk.

Expect requirements to specify records rather than intentions. For most organizations, this level of evidence is contemporaneous, or it simply does not exist.

What leaders need to do
Data practices related to responsible AI have generally included attention to fair representation, the preservation of privacy and high levels of quality to inspire trust. But as organizations expand the use and authority of AI models and agents, new risk factors and legal obligations are coming to light. In response, decision-makers should update their data practices, insisting on tighter control and oversight.

Here’s how:

  1. Move the approval gate upstream. Most governance frameworks concentrate scrutiny at deployment, whereas some of the riskiest and often irreversible decisions were made months earlier, during acquisition.
  2. Give every dataset a named owner and a record. Every data set should have documented sources, collection methods, rights basis, known limitations and permitted uses. All this information should come together as the dataset is being built, not assembled under pressure when a regulator asks.
  3. Treat capture-time privacy as an architecture requirement. If identifying information never enters the data pipeline, most “later problems” never arise. Least-data principles apply here, even as potential benefits of AI might compel the organization to capture as much as possible.
  4. Extend the same questions to what you buy and what you retrieve. Foundation models, third-party datasets and live document indexes all come with their own data history. The standards and requirements you set for your data teams should also apply to third-party due diligence and contracts.

Reiterating the importance of responsible data

If your enterprise is looking to achieve the most with AI, this cannot be done by simply holding the most data. Rather, you should be able to prove, with evidence, where the data came from, on what basis you hold it, and what it is fit for. New models, agents and use cases can move more quickly from proposal to adoption when the organization can quickly answer these questions and count them as settled matters. As with any aspect of responsible AI, good data practices are not just an added layer of control applied at the end of development. They begin much earlier, in decisions that look routine at the time.

Models can be retrained. Data cannot be uncollected.

Share this article

X IconLinked-in Icon

Yann Dietrich

Group Head of Legal - AI, Data & IP

View detailsof Yann Dietrich >
  • Follow Yann Dietrich on LinkedIn

Chris McClean

Global Head of AI Governance and Responsible AI

View detailsof Chris McClean >
  • Follow Chris  McClean on LinkedIn
 

Subscribe for regular insights

Thank you for your interest. You can download the report here.
A member of our team will be in touch with you shortly

Protecting what matters most in the AI economy