The Cost of Unused Data: GDPR, AI, and Europe’s Privacy Bargain
GDPR may be Europe’s most recognisable regulatory export, and one of its most irritating inventions. It changed how companies describe personal data, forced organisations to justify what they collect and gave legal force to the idea that information about people is not free industrial residue. It also became associated with the cookie banner: the repetitive interruption that asks for a choice few people understand and many interfaces are designed to steer.
That association is only partly fair. Most cookie-consent banners arise from the interaction between EU rules governing tracking cookies and GDPR standards for valid consent. Strictly necessary cookies do not require consent, and GDPR permits several legal bases for processing besides consent. Even so, the banner captures a real failure of implementation. A system intended to give people control often reaches them as paperwork, while the organisations processing their data acquire another record showing that a box was clicked.
Artificial intelligence makes the limits of that settlement harder to ignore. Personal data can expose people to surveillance, discrimination, manipulation and administrative power. It can also contribute to medical research, fraud detection, translation, public services and scientific discovery. The relevant question is no longer whether privacy should yield to innovation. That formulation is too crude. Europe needs to decide whether its privacy regime can make legitimate data use sufficiently clear, governed and practical without reopening the extractive economy that made GDPR necessary.
What GDPR Was Built to Protect
A serious criticism of GDPR has to begin with the problem it addressed. Digital systems made personal information unusually easy to collect and combine while leaving the people described by it with little understanding of what was happening. Location histories, purchases, workplace records, social relationships, health information, biometric traces and inferred political preferences could be accumulated simply through ordinary participation in modern life. Data brokers and platforms converted this obscurity into a business model.
Personal data is not politically neutral. It is one of the means by which institutions make people legible. Once a person can be classified through data, they can be priced, targeted, excluded, ranked, monitored or subjected to automated suspicion. The power does not disappear because the underlying record is accurate. Accurate data may make an institution’s intervention more effective while leaving the person affected with no meaningful ability to understand or contest it.
GDPR was Europe’s refusal to treat that information as ownerless exhaust. Its principles—lawfulness, fairness, transparency, purpose limitation, data minimisation, accuracy, security and accountability—require organisations to explain why they process personal data and to remain responsible for what follows. Rights of access, correction, erasure, portability and objection do not equalise the relationship completely, but they make data processing answerable to something beyond the processor’s convenience.
The Regulation itself, however, was not written solely as a shield against use. Its opening article combines the protection of natural persons with rules supporting the free movement of personal data. It permits processing under several lawful bases and gives scientific, historical and statistical research particular treatment when appropriate safeguards are used. Purpose limitation does not mean that information gathered once can never support later learning. It means that reuse must have a lawful and intelligible relationship to the conditions under which the information was obtained.
This distinction matters because the strongest version of the essay’s argument is not that GDPR forgot innovation. The law contains routes for legitimate processing. The harder problem is that those routes can remain difficult to interpret and expensive to operationalise, especially for organisations without specialised legal and technical teams.
When Compliance Becomes Ritual
The cookie banner is a useful emblem because it shows how formal choice can lose much of its substance. The user wants to reach a page. The website wants permission or legal cover. The interface often makes acceptance easier than refusal, and the decision is repeated across hundreds of sites. The resulting click may satisfy part of a compliance process without leaving the person better informed or more capable of controlling subsequent use.
Consent fatigue is not meaningful agency. It is what happens when individual choice is made to carry a governance problem too large for repeated interface decisions. Many data practices cannot be made trustworthy simply by adding another notice. The person asked to agree often lacks the time, expertise and bargaining power needed to evaluate the consequences, while the organisation controls the wording, design and surrounding service.
GDPR is more sophisticated than this experience suggests. Consent is only one lawful basis. Processing may also rest on contract, legal obligation, vital interests, public tasks or legitimate interests, depending on the circumstances. In practice, uncertainty can still push cautious organisations towards visible documentation, repeated approvals and defensive interpretations. Paperwork becomes attractive because it is easier to demonstrate than judgment.
The burden is not evenly distributed. Large companies can employ privacy specialists, build compliance systems and negotiate uncertainty over several years. They also possess extensive first-party datasets that smaller competitors cannot reproduce. A start-up, research group, hospital or municipality may abandon a project before reaching the stage at which a regulator could assess it. The same legal principle can therefore be experienced as a manageable operating cost by one institution and an opaque barrier by another.
It would still be too simple to conclude that GDPR merely protects incumbents. The law also restricts the indiscriminate accumulation of data from which platform power derives, and it gives individuals enforceable rights against the largest processors. Compliance costs can favour scale while substantive constraints limit what scale is allowed to do. The competitive effect depends on enforcement, available guidance, institutional support and whether smaller actors have practical routes towards lawful processing.
Not Every Data Problem Is a GDPR Problem
The phrase “unused data” conceals several different categories. Personal records fall within GDPR when they identify or can reasonably be linked to individuals. Industrial sensor readings, scientific measurements, public statistics and genuinely anonymous datasets may not. Other information may be constrained by copyright, trade secrets, contractual restrictions, confidentiality, professional duties or the simple fact that it is scattered across incompatible systems.
This is important for AI because the training-data problem cannot be reduced to privacy law. A multilingual corpus may contain personal data, but its largest obstacle may be copyright or licensing. Factory data may be largely non-personal but commercially sensitive. Medical information involves both privacy and professional confidentiality. Public-sector records may be lawful to reuse yet remain inaccessible because agencies lack common formats, governance procedures or staff able to prepare them.
Blaming GDPR for every unavailable dataset risks solving the wrong problem. Deregulating personal-data processing would not create clean industrial data, digitise archives, negotiate intellectual-property rights or make incompatible databases interoperable. Europe’s data constraint is legal, technical, commercial and institutional at once.
That does not absolve data-protection law. Artificial intelligence can magnify the consequences of uncertainty because model development often involves large collections assembled for several purposes, indirect acquisition and outputs whose relationship to particular training records is difficult to explain. The usual legal categories remain relevant, but applying them to models is not always straightforward.
The European Data Protection Board’s Opinion 28/2024 on AI models addresses three central questions: when a model may be considered anonymous, when legitimate interest may support development or deployment, and what follows when personal data was processed unlawfully during training. Its answer is not that GDPR prohibits AI training. Legitimate interest may apply in some cases, but the controller must identify a real interest, show that the processing is necessary and balance it against the rights and expectations of affected people.
Anonymity is equally demanding. A model is not anonymous merely because it does not function as a searchable database of names. The assessment must consider whether individuals can be identified from the model or whether their personal data can be extracted with a sufficiently realistic effort. Pseudonymisation reduces risk but does not normally take data outside GDPR if re-identification remains possible.
This case-by-case approach is legally understandable and operationally difficult. It prevents a blanket permission that would legitimise almost any commercial appetite. It also creates uncertainty that organisations with lawyers, compute and patience can navigate more easily than small firms, universities or public-interest projects. The problem is not a simple prohibition. It is the distance between theoretical permission and practical confidence.
The Cost of Non-Use
Privacy debates naturally concentrate on visible misuse: surveillance, discrimination, breach, profiling, manipulation and administrative overreach. The harms of non-use are harder to demonstrate because they concern outcomes that did not occur. A treatment was not discovered, a transport system was not improved, a smaller language was not adequately represented, or a public service remained inefficient. The missing alternative cannot be inspected as easily as a leaked database.
That does not make the cost imaginary. Fragmented medical information can reduce the scale of evidence available for research, particularly for rare conditions. Public administrations that cannot combine records lawfully may fail to detect exclusion or duplicated effort. Smaller language communities may lack the organised corpora needed to build strong translation and speech systems. Industrial firms may collect years of operational information without developing the governance required to use it for maintenance or energy efficiency.
These examples do not create an automatic entitlement to the data. A possible benefit is not enough. Medical information remains intimate, educational predictions can stigmatise children, mobility records can expose private lives, and public-sector analytics can turn suspicion into routine administration. “Social value” is an elastic phrase, and nearly every organisation that wants access can invent a beneficial description of its purpose.
The cost of non-use therefore belongs inside the proportionality analysis rather than above it. A serious decision asks what benefit is plausible, whether the data is necessary, what less intrusive alternatives exist, who bears the risk, how access is controlled and whether affected people can contest the result. Sometimes the answer should be no. Data that remains unused because the proposed purpose is invasive or speculative is not wasted.
The European Health Data Space offers a more useful model than either unrestricted release or institutional paralysis. Its framework for secondary use relies on designated access bodies, applications, data permits and secure processing environments. Researchers do not simply receive a transferable copy of every underlying record. Access can be limited to an authorised purpose and supervised within technical and legal controls.
Whether that system works well will depend on national implementation, administrative speed, data quality and public trust. Its importance lies in the institutional form. It treats privacy and research capacity as a governance problem rather than forcing each hospital, patient and researcher to negotiate the bargain from the beginning.
Europe’s Second Act Has Begun
Europe has not ignored the cost of unusable data. The European Data Union Strategy now explicitly seeks to increase the availability of high-quality data for AI, connect data spaces with AI infrastructure, establish data labs and reduce the cost and uncertainty of accessing datasets. It also proposes greater use of public information, synthetic data and sector-specific arrangements.
This is close to the “second act” that GDPR needs. The first act established limits on extraction and made data processing answerable to rights. The second must build institutions through which lawful use becomes routine enough to compete with both informal scraping and defensive inaction. Data spaces, access bodies and trusted intermediaries matter because they reduce the need for every small organisation to invent its own legal and technical architecture.
The current debate also exposes the danger of moving too far in either direction. The Commission’s broader Digital Omnibus proposal seeks to simplify overlapping digital rules and clarify some aspects of AI development. In their response to the proposal, the European Data Protection Board and European Data Protection Supervisor support simplification but argue that existing GDPR already allows legitimate interest for some AI processing. They also warn that new provisions must not weaken the balancing test, rights of objection or safeguards for sensitive information.
That disagreement is healthier than the older stalemate between “GDPR makes AI impossible” and “any difficulty proves that the organisation should not use the data.” Legal clarity is not the same as permissiveness. A rule can become easier to apply while continuing to prohibit intrusive conduct. Conversely, calling a reform simplification does not guarantee that its effects are narrow or harmless.
The practical goal should be to reduce uncertainty for low-risk and well-governed uses while concentrating regulatory attention on conduct that creates substantial power over people. Secret profiling of vulnerable groups is not morally equivalent to a research consortium analysing protected health data in a controlled environment. A language project working with licensed and filtered material should not need the same governance as a platform combining behavioural traces across billions of users.
Risk-based regulation already claims to recognise such differences. Europe’s challenge is to make the distinctions operational before only the largest institutions can afford to discover what the law permits.
Privacy as Institutional Capacity
A better privacy settlement would not begin by weakening the principle that people deserve protection from institutions that know too much about them. It would reduce the gap between legal permission and practical capability. That requires guidance, standard agreements, shared technical environments and sectoral institutions able to make repeatable decisions.
Trusted research environments can allow analysis without distributing raw records. Federated learning can sometimes train systems across several data holders without pooling every underlying dataset. Differential privacy can reduce the exposure created by aggregate outputs. Secure enclaves, access logs and purpose-bound credentials can make misuse more difficult and more visible. Synthetic data may support some forms of testing where real records are unnecessary.
None of these techniques removes political judgment. Federated systems can still leak information, synthetic datasets can reproduce bias, and secure environments are only as trustworthy as the organisations controlling them. Technical safeguards are useful because they make legal distinctions enforceable, not because they eliminate the need for law.
Europe also needs institutional competence on the side of use. Hospitals, universities, municipalities and small firms require practical support that goes beyond warnings about fines. Templates, specialist access bodies, shared evaluations and regulatory sandboxes can help organisations determine whether a project is lawful before uncertainty has consumed its budget. Enforcement should remain severe where actors conceal extraction or ignore rights, while ordinary procedural errors should not receive more attention than the underlying risk.
This matters for the companion problem of Europe renting its intelligence. Models that reflect European languages, institutions and public priorities require lawful data, compute, technical talent and organisations capable of bringing those elements together. Values written into procurement clauses cannot substitute for systems that can be trained, evaluated and operated under those values.
The cost of unused data is therefore real, but it cannot be calculated by counting everything that has not been fed into a model. Some data should remain inaccessible. Some proposed learning is not worth the intrusion. The failure occurs when a socially valuable and proportionate use is abandoned not because the rights conflict is insoluble, but because Europe has built no practical institution through which it can be resolved.
GDPR’s lasting achievement was to reject the idea that human lives should become raw material simply because technology made extraction possible. The next stage is not to reverse that judgment. It is to build a data economy in which useful access depends on purpose, restraint and accountability rather than on the size of the organisation’s legal department.
Privacy should protect people from what institutions can learn about them. A mature privacy system must also decide, with equal seriousness, what institutions are allowed to learn for them.
Comments
Post a Comment