Promise vs architecture
Every mainstream AI provider now has a privacy story: API inputs aren't used for training by default, retention is limited, enterprise tiers add zero-retention options. Take those commitments seriously — they're contractual, and providers have strong incentives to honor them. But understand what kind of guarantee they are. A policy can change with a terms update. A retention window means your data exists somewhere for that window — subject to breach, insider access, and legal process. In 2025, a US court order in the NYT–OpenAI litigation forced OpenAI to preserve consumer ChatGPT conversations that would otherwise have been deleted — nobody at OpenAI wanted that, and it happened anyway. That's the nature of policy-based privacy: it's real until it collides with something stronger.
Local AI's guarantee is of a different kind. When Llama or Qwen runs on your own GPU via Ollama or LM Studio, the prompt is a buffer in your machine's memory. It isn't transmitted, so it can't be retained; it can't be subpoenaed from a third party, because no third party ever had it; it can't leak in a provider breach, because it isn't in a provider's systems. You can run it with Wi-Fi off and verify. Architecture beats policy not because providers are dishonest, but because architecture doesn't need anyone to keep a promise.
What people actually paste into chatbots
Threat modeling gets concrete fast when you list real usage: medical symptoms and lab results; therapy-adjacent late-night conversations; salary negotiations and resignation letters; client documents you're not authorized to share; code under NDA; immigration and legal questions; your kids' school issues. Each is innocuous until it's aggregated into the most detailed behavioral profile any company has ever held about you — more candid than search history, because people talk to chatbots like confidants.
Cloud tiers differ meaningfully here, and honesty requires saying so: API access with a business agreement is far more private than consumer chat apps (no training use, shorter retention, no human review by default), and enterprise tiers are better still. If you stay in the cloud, moving your sensitive use from a consumer app to the API tier is the single biggest privacy upgrade available. The free-tier directory notes which providers offer it.
What privacy costs you in quality
The honest trade: no local model matches GPT-4o or Claude on hard reasoning, obscure knowledge, or polish. The gap has narrowed dramatically — Qwen 3 14B on a 16 GB GPU or a Mac's unified memory handles everyday drafting, summarizing, and Q&A at a level indistinguishable from cloud for most prompts — but it exists, and pretending otherwise sells the switch dishonestly. The pragmatic pattern most privacy-conscious users land on is a split stack: local by default (especially anything personal), cloud API for the occasional task that genuinely needs frontier quality, with the sensitive details stripped. Privacy isn't all-or-nothing; routing 90% of your tokens locally is 90% of the win.
The hardware is the easy part
A Mac mini M4 Pro (~$1,399) runs 8–14B models silently on your desk at ~65 tokens/sec; any PC with a 12–16 GB GPU does the same — check what your machine runs before spending anything. Cost-wise, against Claude Sonnet pricing at 500k tokens/day, the Mac pays for itself in about 15 months; against mini-class cloud tiers it never does (they're a few dollars a month, as our budget comparison shows). You're not buying it to save money. You're buying the only version of AI where nobody else is in the room.