Personal Intelligen

Privacy Risks of Consumer AI Apps

Popular AI apps collect far more data than users realize, and share it in ways they cannot control.

Columnist · · 10 min read
Cover illustration for “Privacy Risks of Consumer AI Apps”
AI Consumer MOdels · July 26, 2026 · 10 min read · 2,180 words

Start with what's actually being gathered, because most people would guess wrong. A 2025 Surfshark study tracked 35 data types across ten popular AI chatbots. The collection goes well past conversation content: device identifiers, usage patterns, location signals, and in some cases biometric data. Meta AI and Google Gemini collected the most categories. Most users suspect something is off; a 2025 survey found 81% believe AI companies collect more than necessary. But suspecting excess and knowing the specifics are meaningfully different things. That raises an important question: if most users already sense that something is wrong, why does the gap between suspicion and actual knowledge persist?

The third-party sharing dimension is where this gets uncomfortable in a specific, concrete way. Late 2025: Meta announced it would use data from your conversations with its AI to personalize content and ads. No opt-out. The categories they explicitly carved out from this practice are telling: religious views, sexual orientation, health conditions, political views, racial and ethnic origin. They drew those lines because the data was sensitive enough to demand them. That's an implicit acknowledgment of the problem, not a reassurance about it.

OpenAI, same year, began testing conversation-topic-targeted advertising in ChatGPT for free and lower-tier users in the US. Stanford HAI's congressional testimony in November 2025 named the dynamic plainly: large platforms are actively working out how to monetize chatbot-derived user data across their broader businesses.

So the app you're talking to is rarely the only entity receiving what you share. Advertising ecosystems, cloud infrastructure providers, parent companies, they're all downstream. NowSecure found that DeepSeek transmits data to Volcengine, ByteDance's cloud platform. That's an infrastructure-level data flow, invisible to the user, that no amount of careful typing prevents. You can choose your words deliberately and still have no meaningful control over where they end up.

How long apps keep your data, and what they do with it to train their models

Retention policies, when you actually read them, are alarming in specific ways. Anthropic's Claude changed its terms in 2025: user chats can now be retained for up to five years if users permit training. The opt-out exists, buried in settings most users navigate to. A 2025 Stanford HAI analysis of privacy policies from six major chatbot developers found that many collect and use personal information disclosed in chats by default to train their systems, including health data, biometric data, and in some cases children's chat data, some retaining it indefinitely.

The accompanying congressional testimony was direct: developers are incorporating chatbot-derived user data into model training without meaningful oversight, and their privacy policies show a lack of transparency about whether they take steps to mitigate privacy risks.

Training data extraction is not theoretical. Research presented at two major academic conferences in 2025 demonstrated scalable extraction of training data from production language models. Personal data shared in a conversation can, under certain conditions, be recovered from model outputs later. Something you typed months ago can surface in a response to a completely different user asking a completely different question. But how does this affect our original promise? The promise of a private, contained conversation was not as solid as it felt.

Here's the asymmetry worth sitting with. You have a narrow, time-limited interaction with a chatbot. The data that interaction generates persists in model weights or server logs for years, doing things you never anticipated and cannot observe. The interaction feels bounded. The record it creates is not.

The inference problem: what AI can deduce from data you didn't knowingly share

This is the part that offers no clean solution, even for careful users. AI systems can infer sensitive attributes, including health conditions, sexual orientation, race, and religious affiliation, from pieces of information that seem unremarkable in isolation. You don't have to disclose a diagnosis. You just have to ask a few questions that happen to correlate with one. But what if the most revealing data you share isn't the information you're consciously disclosing — what if it's the questions you think are harmless?

What makes this structurally different from ordinary privacy risk is that the inference process is not fully understood even by the engineers who built the model. That makes it difficult for users to anticipate what will be deduced from what. You cannot consent to inferences you cannot predict.

Re-identification risk compounds this further. A University of Chicago Medicine patient sued Google after medical records used to train AI models were alleged to contain re-identifiable information, despite claimed anonymization. Anonymization sufficient for traditional databases doesn't hold once AI systems can cross-reference multiple weak signals to reconstruct an identity. The old standard, the one regulators and companies both relied on, no longer applies in the same way.

The profiles that result from all of this feed ad targeting, pricing decisions, or other purposes entirely opaque to users. There's no moment where the feedback loop becomes visible enough to object to. That invisibility is not incidental; it's baked into how these systems operate.

Security failures that have exposed user data at scale

Inference risks assume the data stays within the company that collected it. Security failures remove even that constraint, and recent history offers a fairly grim catalog.

February 2026: a security researcher accessed 300 million messages from over 25 million users of Chat & Ask AI through an exposed Firebase database. The messages included discussions of illegal activity and requests for suicide assistance. Root cause: Security Rules left public. A well-known, entirely preventable error. A broader scan of 200 iOS apps found 103 had the same vulnerability, collectively exposing tens of millions of stored files.

January 2025: Wiz Research found a publicly accessible misconfigured database belonging to DeepSeek containing more than one million log entries, including plaintext chat history and API authentication keys, requiring no authentication to access. Separately, hard-coded encryption keys and transmission of unencrypted user and device data. And separately still: under China's 2017 National Intelligence Law, DeepSeek is legally required to cooperate with Chinese intelligence agencies upon request. That's a legal mechanism with no equivalent in Western jurisdictions, running in parallel with the technical vulnerabilities.

February 2026: Video AI Art Generator and IDMerit leaked over 12 terabytes of data, including identity-verification documents from users across approximately 25 countries, through misconfigured Google Cloud Storage. September 2025: Trend Micro found that Wondershare RepairIt contradicted its own privacy policy by collecting and retaining sensitive user photos through overly permissive cloud access tokens embedded in the app's code.

The pattern across these incidents is consistent and, frankly, unglamorous. Cloud misconfiguration is the dominant failure mode. None of these were sophisticated attacks. These were basic infrastructure errors at companies that had collected far more data than they were prepared to secure. The AI was sophisticated. The security infrastructure protecting it was not.

Children's exposure to AI data collection, and why it's a distinct problem

Three out of four teenagers use AI chatbots. One in three reports being uncomfortable with something a chatbot said to them. The data collection and training practices that apply to adults apply to minors too, frequently with fewer guardrails, and what minors are sharing tends to be more emotionally raw.

It is also worth considering companion bots as a distinct category. In 2025, Google rolled out Gemini targeted to children under 13. EPIC characterized it as having insufficient safeguards. Stanford HAI's analysis found that some companies collect and train on children's chat data, including sensitive categories, retaining some of it indefinitely. Existing child-protection laws were written for a different technological era, one where the data being collected was far more legible and far less intimate than what a teenager shares with a companion AI at 11pm.

The regulatory response has been real. An updated COPPA rule effective June 2025 established that using a child's personal information to train AI requires separate, verifiable parental consent. Biometric identifiers, including voiceprints and facial templates, were added to the definition of personal information, and indefinite retention of children's data was prohibited. The FTC settled with Disney for $10 million in September 2025 over collection of children's data for targeted advertising without parental consent. Italian regulators fined an AI companion chatbot operator €5 million in May 2025 for privacy violations.

These numbers describe enforcement after harm. The structural problem remains: children are more likely to share emotionally sensitive content with companion AI, less likely to read privacy policies, and less equipped to assess what long-term disclosure costs them. The data they share today gets retained and trained on for years. The conversation ends. The record does not.

The regulatory landscape users are now implicitly relying on

Most people opening a chatbot aren't thinking about regulatory frameworks. But those frameworks are operating quietly in the background, and knowing their actual shape changes how much weight you should put on them.

In the United States, the landscape is fragmented in ways that matter practically. There is no federal AI privacy law. The FTC has warned that deploying AI tools without proper consent constitutes unfair and deceptive practice, but enforcement remains selective and retrospective. State law is filling gaps unevenly: Utah's AI Policy Act in 2024 was the first major state statute specifically governing AI; Colorado, Connecticut, and Texas passed AI transparency laws in 2024 and 2025 covering automated decision-making. State attorneys general intensified AI-related consumer-protection scrutiny throughout 2025, which is meaningful protection, though its value depends heavily on which state you live in.

The European Union offers more comprehensive architecture. The EU AI Act's prohibited practices took effect February 2025; rules governing General-Purpose AI models became effective August 2025; full applicability arrives August 2026. The highest fine tier reaches €35 million or 7% of global annual turnover, a ceiling higher than GDPR. That reflects a deliberate legislative judgment that the most serious AI risks exceed traditional data-protection violations in kind. Total GDPR fines across all sectors exceeded €5.65 billion as of spring 2025.

Litigation is filling additional gaps. Plaintiffs have begun applying wiretapping statutes to AI-powered conversational analytics, alleging that AI vendors used customer communications for model training without consent. Illinois's Biometric Information Privacy Act remains significant: it requires notice, written consent, and retention limits before collecting biometric identifiers, and AI dramatically increases exposure because collection can occur continuously and at scale. Meta's $650 million BIPA settlement in 2021 for collecting facial geometry without consent remains the largest single AI privacy settlement on record. LinkedIn received a €310 million GDPR fine in 2024 for consent manipulation in AI-driven behavioral profiling.

Regulatory protection exists. It is inconsistent, jurisdiction-dependent, and largely reactive by design, meaning it follows harms already incurred rather than harms in progress. One might argue that this is simply how legal frameworks respond to fast-moving technology — and that argument has merit. But it does mean you can't outsource the judgment entirely to regulators and assume the gap is covered.

What users can actually do to reduce their exposure

Avoiding AI tools entirely is both impractical and an overcorrection. The more useful standard: use them with the same deliberate judgment you'd apply to any service that holds sensitive personal information.

Before using a new AI app, determine whether it's a standalone product or an embedded feature within a larger platform. Data flows differ substantially depending on the answer, and knowing which one you're dealing with tells you where to look for the relevant terms. Retention and training opt-out settings exist, often buried, but they're there. Anthropic's five-year default retention can be reduced by opting out, if you know to look for it.

For apps requesting access to photos, contacts, or files, ask whether that access is actually necessary for the stated function. If the answer isn't immediately obvious, treat that ambiguity as information worth acting on.

Some categories warrant particular caution regardless of platform: medical details, financial account information, anything involving children's identities. These appear repeatedly in the most serious breach and enforcement cases. Workplace information is a specific concern worth naming separately; research from Harmonic Security found that enterprises upload roughly 1.3 gigabytes of files to generative AI tools every quarter, with about 20% containing sensitive data. Personal use carries equivalent risk, without the institutional security layer enterprises sometimes provide.

For parents, the updated COPPA rule creates new rights around children's AI data, including the right to separately consent to training use. Those rights only matter if parents know they exist. Companion AI apps warrant at least the same scrutiny as social media, probably more, given what teenagers actually share in those conversations.

Worth being clear about what users cannot control. Infrastructure security failures, the misconfigured databases behind the Chat & Ask AI and DeepSeek breaches, are entirely outside user control. Those users had done nothing wrong. Inference risks mean that careful, limited disclosure doesn't fully bound what a platform ultimately knows about you. Some exposure is structural, not behavioral, and no amount of thoughtful typing eliminates it.

What you can control is narrower: choosing platforms with transparent retention policies, using opt-out mechanisms when they exist, and calibrating disclosure the same way you would with any record that will outlast the conversation that created it.

Sources

  1. statista.com
  2. wilmerhale.com
  3. krebsonsecurity.com

More in AI Consumer MOdels