Stop doing cost math; it's not about cost
Run the numbers once and move on: at a working writer's volume (~300k tokens/day of drafting and feedback), GPT-4o costs about $43/month, and mini-class models cost pocket change. A $1,399 Mac mini M4 Pro takes almost three years to break even against that — and never against the budget tiers. Anyone selling you local AI for writing on economics is doing the arithmetic wrong. The decision lives elsewhere.
Where it actually lives: what you're writing
A manuscript is different from a code snippet. Unpublished fiction, a memoir chapter, a client's ghost-written book, a sensitive investigation, therapy-adjacent journaling — this is material where "processed under provider terms with a retention window" lands differently than it does for boilerplate code. Every draft you paste into a cloud assistant exists, for some window, on infrastructure you don't control, subject to breach, legal process, and policy evolution (the 2025 litigation-hold order that froze deleted consumer ChatGPT conversations is the canonical example). A local model on your own machine offers a categorically different promise: the unsent sentence stays unsent. For ghostwriters and journalists with source-protection obligations, that's not a preference — it's professional hygiene. The privacy comparison treats this in depth.
The quality question, without cope
Honesty cuts the other way on quality. For sentence-level work — rephrasing, tightening, tone shifts, "give me five alternatives to this clunky line" — local 14–30B models (Qwen 3 14B on a MacBook, Qwen 3 32B or Gemma-class on a 24 GB GPU) are excellent, and most writers cannot reliably distinguish their line edits from a frontier model's. For structural work — "what's wrong with this chapter," developmental feedback across 40,000 words, keeping a book's argument coherent — frontier models are visibly better and the gap has not closed. They hold more in working memory, and their feedback reads like a good editor rather than an eager workshop peer.
So the split, again, follows the work: local for the hundred daily micro-interactions of drafting; a cloud session (with material you're comfortable sharing, or suitably excerpted) for the periodic structural pass.
The consistency bonus nobody advertises
A quiet local advantage for professionals: the model never changes under you. Cloud models get silently updated; a prompt that produced your voice in March produces something subtly different in June. A local model is a frozen artifact — the same weights, the same temperament, for as long as you keep the file. Writers who've built elaborate style prompts learn to treasure this. (Corollary: you also don't get free upgrades. You choose when to move.)
What to run, on what
Writing is the least demanding mainstream LLM workload — no giant contexts, no tool calls, modest speed needs (you read slower than any model generates). A Mac mini M4 Pro (24 GB, $1,399) or any 12–16 GB GPU runs Qwen 3 14B-class models silently and instantly; a used $500-class build does the job too. LM Studio is the writer-friendly on-ramp: a chat GUI, model browser, no terminal. Try it against your actual work-in-progress for a week before deciding anything — what your current machine runs is free to check.