Artificial Intelligence and the End of Informational Privacy

My illustration entitled: “The Privacy Eclipse” – a radiant AI sphere passes before a symbolic moon marked “PRIVACY,” casting the city into an informational eclipse. I protect the citizens beneath a translucent dome of encryption and decentralized identity.


Privacy once depended upon a comparatively simple boundary: information was either revealed or it was not.

A person disclosed a name, address, photograph or medical condition, and another party received that information. If the fact was never communicated, it appeared to remain private.

Artificial intelligence is dissolving this boundary.

An AI system does not need to receive an intimate fact directly if it can infer that fact from patterns scattered across ordinary data. Purchases, movements, search queries, writing style, social connections, viewing habits and device activity can be combined to predict characteristics that the individual never consciously disclosed.

The distinction between known and unknown is consequently being replaced by a distinction between what has already been inferred and what remains computationally inferable.

This is not simply a larger version of conventional data collection. It is a transformation in the nature of informational power.

Artificial intelligence does not merely remember what a person revealed. It can derive what the person attempted to keep private.

The Old Model of Informational Privacy

Informational privacy has traditionally concerned the collection, storage, use and disclosure of information about an individual. The person provides data to a government, employer, bank, hospital or company, and rules determine what that institution may do with it.

This model remains important. Organizations should still collect no more than necessary, secure what they retain and refrain from using information for incompatible purposes.

But the model assumes that sensitive information enters the system as sensitive information. A medical condition appears in a medical record. Political affiliation appears in a membership database. Religious belief appears in an explicit declaration.

Machine learning disrupts this assumption. Sensitive knowledge can emerge from information that appears harmless when viewed separately. The private fact may never have been present in the original dataset as a field. It can be created later as an inference.

The institution no longer needs to ask the person the intimate question. It can ask the model.

Four Forms of Personal Information

To understand the change, it is useful to distinguish four forms of personal information:

  1. Provided information is deliberately supplied by the individual, such as a name entered into a form or a photograph uploaded to a service.
  2. Observed information is recorded through behavior, including location, browsing activity, purchases, communication patterns and device interactions.
  3. Derived information is calculated from existing records, such as total spending, average travel distance or the number of interactions with a particular person.
  4. Inferred information is a probabilistic conclusion about a characteristic, intention or future action that was not directly supplied.

Existing privacy controls often concentrate upon the first category. They ask whether the person consented to provide information and whether the organization was authorized to store it.

Artificial intelligence derives its greatest power from the remaining categories. A person may refuse to disclose a sensitive fact while continuing to generate behavioral signals from which the same fact can be predicted.

The information was not surrendered as a statement. It was reconstructed as a probability.

From Metadata to Inference

In Metadata Is Power, I argued that the circumstances surrounding communication can reveal identities, locations, associations and routines even when the content remains protected.

Artificial intelligence magnifies this power. It can compare a person’s metadata with patterns derived from millions of other people. A small signal that means little in isolation may become predictive when placed within a population-scale model.

The frequency of late-night activity may contribute to a health inference. Changes in travel or purchasing behavior may suggest a life event. A network of contacts may reveal political, religious or professional associations. Language patterns may influence judgments about personality, education or emotional condition.

The system may not know these things with certainty. It may not need to. A sufficiently confident prediction can alter which advertisement, price, opportunity, investigation or restriction a person receives.

In digital governance, probability can exercise power before certainty has been established.

The Inference Gap

An inference gap appears when an institution possesses knowledge about a person that the person did not knowingly disclose and may not know has been created.

On one side of the gap is the individual, who sees isolated actions: a purchase, a search, a journey or a message. On the other is the institution, which sees those actions as features within a predictive model.

The individual does not know which signals matter, which datasets have been combined, what conclusions have been drawn or how those conclusions influence subsequent decisions.

This asymmetry is a new form of informational power. The institution can interpret the individual while remaining difficult for the individual to interpret.

The danger exists whether the inference is accurate or false. An accurate inference can invade privacy. A false inference can produce discrimination, suspicion or exclusion. In both cases, the person may have no practical means of discovering or contesting the conclusion.

Inference Privacy

Digital civilization therefore requires a concept broader than conventional data confidentiality. I call this Inference Privacy.

Inference Privacy is the capacity of an individual to prevent, limit, understand and contest the computational derivation of sensitive conclusions from personal or behavioral data.

Inference Privacy does not establish that no one may ever reason about another person. Human beings make observations and draw conclusions as part of social life. Nor does it prohibit every beneficial use of predictive analysis in medicine, security or scientific research.

It establishes that automated inference becomes a matter of privacy when it is conducted systematically, at scale and with consequences for the individual.

A model that predicts a serious medical condition, political preference, psychological vulnerability or financial desperation exercises informational power even if its input data was lawfully collected. The legitimacy of the original collection does not automatically legitimize every future inference.

The Reconstruction of Identity

Removing names from a dataset does not necessarily make the people represented within it anonymous.

In 2008, Arvind Narayanan and Vitaly Shmatikov demonstrated that supposedly anonymous movie-rating records could be connected with publicly available information to identify users. Their research showed how combinations of sparse behavioral data can function as fingerprints.

Artificial intelligence expands the range of patterns that can be used for such reconstruction. Location sequences, writing styles, purchasing histories, social graphs and device characteristics can distinguish one person from a population even when direct identifiers have been removed.

Anonymization is consequently not a one-time procedure. Its effectiveness depends upon what other information exists and what computational methods become available.

A dataset that appears anonymous today may become identifiable tomorrow when combined with new records or analyzed by a more capable model. Privacy risk persists across time because the data remains while computational power advances.

Ordinary Data Can Reveal Extraordinary Facts

The privacy threat does not begin only with medical records, biometric databases or confidential messages. It begins with ordinary digital traces.

In 2013, researchers Michal Kosinski, David Stillwell and Thore Graepel demonstrated that private traits could be predicted from digital records of behavior. Their study used Facebook “Likes” to estimate characteristics that were not necessarily explicit in the individual records.

The broader lesson is more important than any particular prediction. Information does not possess a fixed degree of sensitivity. Its sensitivity changes according to the analytical capabilities applied to it and the other information with which it can be combined.

A music preference may contribute to a personality profile. A shopping pattern may contribute to a health prediction. A sequence of locations may reveal religious attendance, political activity or an intimate relationship.

There may be no such thing as permanently harmless personal data. There may only be data whose inferential value has not yet been discovered.

Prediction Before Action

Traditional surveillance observes what a person has done. Predictive systems attempt to estimate what a person may do next.

This changes the temporal structure of power. An individual can be classified before applying for a job, investigated before committing an offense, offered unfavorable terms before demonstrating financial behavior or manipulated before expressing an intention.

The person is governed not only through his or her history but through a machine-generated future.

Predictions may become self-reinforcing. If a model classifies a person as risky, the individual may receive fewer opportunities. The resulting lack of opportunity can then be recorded as additional evidence of risk. Classification shapes the conditions from which future classifications are made.

The model does not merely describe reality. When connected to institutional decisions, it participates in constructing reality.

Consent Cannot Govern Unknown Inferences

Many privacy systems rely upon consent. An organization presents terms, the individual agrees and information is collected.

This approach becomes inadequate when neither party can predict every conclusion that future models may derive from the data.

A person cannot meaningfully consent to an inference that has not been described, an analytical method that does not yet exist or a purpose that will be invented years later. A broad clause permitting “improvement of services” cannot legitimately authorize every possible conclusion about a person’s health, beliefs, relationships or vulnerabilities.

Consent also weakens when participation is unavoidable. Refusing data collection may mean losing access to employment, banking, education, communication or public services.

The legal appearance of choice should not conceal architectural dependence.

Inference Privacy therefore requires more than agreement at the moment of collection. It requires continuing limits upon what may be derived, for what purpose and with what consequences.

Generative AI and the Memory of the Model

The public release of increasingly capable generative systems has made artificial intelligence visible to a much wider population. The introduction of ChatGPT in November 2022 demonstrated how a model trained on enormous quantities of text can generate fluent responses across many subjects.

Generative models create a distinct privacy question. Training data is not ordinarily stored as a simple searchable archive inside a model. Yet research has shown that large language models can memorize portions of their training data and may sometimes reproduce identifiable information.

In 2021, Nicholas Carlini and his co-authors demonstrated methods for extracting examples of training data from a large language model. The study established that model access can sometimes reveal information incorporated during training.

This complicates familiar ideas of storage and deletion. Removing a document from a database is conceptually straightforward. Determining whether information influenced a trained model, how it affected the model and whether it can later be reproduced is far more difficult.

A model can retain the influence of data without functioning like a conventional record system. Privacy protection must therefore extend across the entire AI lifecycle: collection, preparation, training, deployment, querying, updating and retirement.

The Three AI Privacy Threats

Artificial intelligence creates at least three distinct privacy threats.

1. Memorization

A model may retain and reproduce information contained within its training data. This risk is particularly serious when datasets contain personal, confidential or improperly collected material.

2. Re-identification

A model may connect anonymous or pseudonymous records with external information and reconstruct the person represented by them.

3. Inference

A model may derive sensitive characteristics that were never explicitly present in the data supplied by the individual.

These threats require different responses. Encryption can protect stored data against unauthorized access but does not prevent an authorized system from generating invasive inferences. Anonymization may remove obvious identifiers but not prevent re-identification. Deleting a source record may not remove its influence from an already trained model.

Privacy must therefore be designed around the capability of the entire system, not merely the confidentiality of the original database.

The Limits of Encryption

Cryptography remains essential. Encryption can protect information while it is transmitted or stored, preventing unauthorized parties from reading it. Strong authentication and cryptographic access controls can reduce breaches and limit who enters a system.

But encryption alone cannot solve inference privacy.

If an AI provider is authorized to decrypt the data and analyze it, the provider may still derive sensitive conclusions. The information was protected from outsiders but not from the institution performing the computation.

This distinction is crucial. Confidentiality asks who can read the input. Inference Privacy asks what the authorized reader can learn from it.

Cryptography must therefore be combined with data minimization, local processing, purpose limitations and privacy-preserving computation.

Privacy-Preserving Artificial Intelligence

Artificial intelligence does not have to be constructed as a centralized system of informational extraction. Several technical approaches can reduce privacy risk.

On-Device and Local Computation

When computation occurs on a person’s own device, raw information need not always be transferred to a remote provider. A model may analyze photographs, messages or behavioral information locally and disclose only the result chosen by the user.

Local computation is not automatically private. Applications can still transmit data, and device software may remain controlled by an external platform. Nevertheless, processing close to the individual can reduce unnecessary central collection.

Federated Learning

Federated learning allows a model to be trained across multiple devices or institutions without gathering all raw training data into one central repository. Participants contribute model updates rather than transferring complete local datasets.

This reduces some risks but does not eliminate them. Model updates can themselves leak information, and the coordinating system may remain centralized. Federated learning must be combined with secure aggregation and other safeguards.

Differential Privacy

Differential privacy introduces carefully calibrated randomness so that the inclusion or exclusion of one person’s data has a limited effect upon the released result. It can allow useful population-level analysis while reducing what can be learned about a particular individual.

Its effectiveness depends upon implementation choices, including the privacy budget and the number of analyses performed. The label alone does not guarantee protection.

Secure Multiparty Computation

Secure multiparty computation enables participants to calculate a result jointly without each participant revealing all of its private inputs to the others. This can support cooperation among institutions that need a shared answer but do not need access to one another’s complete datasets.

Homomorphic Encryption

Homomorphic encryption allows certain computations to be performed upon encrypted data. The party performing the calculation may produce an encrypted result without receiving the underlying information in readable form.

These techniques remain computationally demanding and cannot yet replace every conventional AI architecture. Their importance lies in demonstrating that centralized exposure is a design choice, not an unavoidable law of computation.

Data Minimization Must Include Inferences

Data minimization is often interpreted as collecting fewer database fields. AI requires a broader interpretation.

An organization should ask not only what data it stores but what knowledge its systems can generate. A limited set of inputs may still support extraordinarily sensitive conclusions.

Inference-aware minimization requires several questions:

  • What predictions can be derived from the collected information?
  • Are those predictions necessary for the stated service?
  • Could the same objective be achieved with local or less granular processing?
  • Are inferred profiles retained after the immediate purpose has ended?
  • Can the inference be used in unrelated institutional decisions?
  • Can the individual discover and challenge it?

A system has not minimized data if it collects fewer facts while maximizing the conclusions extracted from them.

The Right to Contest an Inference

A person may correct a false address in a database. Correcting an inference is more difficult because the system may treat it as a probabilistic output rather than a factual record.

The organization may claim that the model is proprietary, too complex to explain or accurate on average. None of these claims resolves the individual harm caused by a consequential classification.

When an inference affects employment, credit, insurance, healthcare, policing, education or access to an essential service, the person should be able to know that an automated assessment occurred, understand the principal basis of the decision and present evidence against it.

Contestability must be capable of changing the result. A human reviewer who merely repeats the model’s output does not provide meaningful review.

The European Union’s General Data Protection Regulation establishes protections concerning profiling and decisions based solely upon automated processing in certain circumstances. The underlying principle should extend further: consequential computational judgments must not become unchallengeable authority.

Artificial Intelligence and the Architecture of Power

AI is often described as a tool. Yet a tool that observes populations, predicts behavior and influences access to opportunities also becomes an institution of power.

The model’s owner determines what data enters the system, what objective is optimized, which errors are accepted and who receives the results. The interface determines what the individual can see. The database determines what can be remembered. The permissions determine who can act upon an inference.

This extends the theory developed in The Architecture of Power. Artificial intelligence adds an interpretive layer to technological architecture. The system no longer merely stores information or enforces explicit rules. It generates judgments.

Whoever controls that interpretive layer can influence how the individual becomes visible to institutions.


My illustration “The Privacy Eclipse” work-in-progress. The art represents that unrestrained AI could obscure the boundary between public information and private life.


The Danger of Centralized AI

The most capable AI systems require significant data, computing infrastructure and specialized knowledge. These requirements encourage concentration among governments and large corporations.

Centralization can produce technical benefits. Shared infrastructure can support security, research and services too expensive for individuals to build independently. Cypherpunkism does not reject every centralized system.

But concentration becomes dangerous when a small number of institutions possess both population-scale data and the models capable of interpreting it. The institution that controls the data can improve the model, and the institution that controls the model can extract greater value from the data. Each advantage strengthens the other.

This cycle resembles the dependency examined in Cypherpunkism Against Digital Colonialism. Populations may supply the language, images, behavior and knowledge used to train AI systems while possessing little influence over the resulting models or the value they generate.

Decentralization, open standards and privacy-preserving computation can provide checks upon this concentration. But openness must be balanced with the obligation not to expose personal training data or create systems whose risks cannot be responsibly governed.

Twelve Principles for AI and Inference Privacy

A Cypherpunkist approach to artificial intelligence should include at least twelve principles:

  1. Defined purpose: Personal data should be used to train or operate an AI system only for clear and legitimate purposes.
  2. Inference limitation: Authorization to collect data should not automatically authorize every sensitive conclusion that can be derived from it.
  3. Data minimization: Systems should collect the least information required and avoid retaining raw data without necessity.
  4. Local processing: Computation should occur on user-controlled devices when central transfer is unnecessary.
  5. Cryptographic protection: Information should be protected during storage, transmission and, where feasible, computation.
  6. Unlinkability: Separate contexts should not be combined into universal profiles without legitimate justification.
  7. Transparency: People should know when AI is generating consequential inferences about them.
  8. Contestability: Individuals should be able to challenge inaccurate, disproportionate or improperly used conclusions.
  9. Retention limits: Raw data and inferred profiles should expire when their legitimate purpose has ended.
  10. Independent assessment: High-impact systems should be available for appropriate auditing of privacy, security, discrimination and misuse.
  11. Architectural alternatives: No unnecessary single provider should control the data, model, identity and decision channel simultaneously.
  12. Human responsibility: Institutions must remain accountable for decisions made with AI and may not transfer moral responsibility to an algorithm.

These principles are consistent with the OECD AI Principles, adopted in 2019, and the UNESCO Recommendation on the Ethics of Artificial Intelligence, adopted in 2021. Both recognize that trustworthy AI must respect human rights, privacy, transparency and accountability.

Applying the Cypherpunkist Test to AI

The Cypherpunkist Test asks who a technology ultimately empowers. Applied to an AI system, it requires the following questions:

  • Who supplies the data?
  • Who controls the model?
  • What private characteristics can the model infer?
  • Does the person know that the inference exists?
  • Can data from separate contexts be combined?
  • Can the system operate locally rather than transferring raw information?
  • Can training data be extracted or memorized?
  • Who receives the prediction?
  • What decision is made because of it?
  • Can the individual challenge the inference?
  • Can the individual withdraw data or leave the system?
  • Who is accountable when the inference causes harm?

An AI system is not sovereignty-enhancing merely because it is intelligent, efficient or personalized. It is sovereignty-enhancing when its capabilities remain compatible with the individual’s authority over personal information and consequential decisions.

The End of Privacy—or the End of an Old Definition?

Artificial intelligence does not make every form of privacy impossible. The title of this essay does not announce that resistance is futile or that informational boundaries have already disappeared.

It announces the end of an inadequate definition.

Privacy can no longer mean only control over facts consciously disclosed. It must include control over observation, aggregation, correlation, prediction and inference.

Protecting a person’s name while exposing a behavioral fingerprint is not privacy. Removing an identity field while preserving enough information to reconstruct the individual is not anonymity. Obtaining consent for data collection while concealing the conclusions later derived from it is not informational self-determination.

The privacy question of the AI age is no longer simply:

What does the system know about me?

It must also become:

What can the system infer about me, and what power does that inference give it?

Conclusion: The Right Not to Be Computed Into Submission

Artificial intelligence can discover medical patterns, improve accessibility, translate languages, assist scientific research and expand human knowledge. Its ability to find relationships invisible to ordinary observation is precisely what makes it valuable.

The same ability makes AI a profound privacy technology—either a technology for protecting human boundaries or a technology for dissolving them.

A society of constant inference would not require every person to confess. Systems could estimate beliefs, vulnerabilities and intentions continuously, then shape opportunities before individuals understood how they had been classified.

Digital Sovereignty requires more than secrecy in such a world. It requires the capacity to limit computational observation, understand consequential inferences, challenge automated judgments and prevent one institution from becoming the permanent interpreter of human life.

The individual must not become transparent while the model remains opaque.

The right to privacy must include the right not to have every ordinary action transformed into an intimate prediction.

Privacy is sovereignty.

Cryptography is applied freedom.

Decentralization is a check on power.

Code is political architecture.

Digital sovereignty belongs to the individual.


References