Your Home GPU in Your Pocket: Run LM Studio Models From Your Phone With LM Link

LM Studio's new LM Link and the Locally iPhone app let you run your biggest local models — even 70B+ — from your phone. The model stays on your desktop; only the chat travels, end-to-end encrypted. Here's how it works and why it matters.

June 27, 20269 min readJakub Rusinowski

There are two completely different ways to "run AI on your phone," and people constantly confuse them.

The first is on-device inference: a small 1B–4B model running directly on your phone's chip, fully offline. We covered that in Can You Run a Local LLM on Your Phone? — it's great for airplane mode, but you're capped at tiny models.

The second is what LM Studio shipped in June 2026, and it changes the math entirely. With LM Link and the Locally iPhone app, your phone becomes a remote control for the powerful machine you already own. The model runs on your desktop's GPU. Your phone just sends the prompt and streams back the answer — over an end-to-end encrypted connection, with nothing routed through a cloud AI provider.

That means the 70B model that would never fit on a phone is suddenly usable from your phone, anywhere you have a connection.

It's worth being precise, because the names get mixed up:

  • Locally is the iPhone and iPad app you install from the App Store. It's the chat interface in your pocket.
  • LM Link is the remote-access layer that connects that app (or other tooling) to an LM Studio instance running somewhere else — your desktop, your home server, your studio Mac.

You need both: LM Studio on a host machine doing the actual inference, and the Locally app on your phone talking to it through LM Link.

> Note (June 2026): At launch, LM Studio 0.4.16 gated LM Link behind a request-only Preview. A follow-up build on June 8, 2026 removed the waitlist, so it's open to everyone. It's free during the Preview, with free and paid tiers planned for general availability. Locally is iPhone/iPad only for now — Android hasn't been announced.

How It Actually Works (The Part That Matters for Privacy)

This is the key design decision, and it's a genuinely good one: LM Link runs on top of a custom Tailscale mesh VPN.

Here's what that buys you:

  • Inference runs on your Mac, PC, or Linux box. The model weights never leave the host. Your phone is a thin client.
  • All communication between your devices is always end-to-end encrypted, thanks to the underlying Tailscale mesh.
  • Your devices are never exposed to the public internet. There's no open port to scan, no reverse proxy to harden, no "I accidentally published my LLM to the world" mistake. The connection is a private peer-to-peer tunnel between machines you own.
  • Chats stay on your devices. According to LM Studio, the only thing that ever touches their servers is a device discovery list — the metadata needed to find your other devices. Your prompts and responses don't.

Both devices sign into the same LM Studio account, and that shared identity is what authenticates the link. No QR-code dance with IP addresses, no port forwarding, no dynamic DNS.

If you've ever tried to expose a local LLM server to your phone the manual way — binding to 0.0.0.0, opening a firewall port, fronting it with a reverse proxy and auth, maybe running ngrok — you know how fiddly and risky that is. LM Link automates the safe version of all of it.

Setting It Up: Three Steps

The whole thing is about a five-minute setup once LM Studio is installed.

1. On your desktop (the machine with the GPU or Apple Silicon):
  • Update LM Studio to 0.4.16 or later.
  • Sign in to your LM Studio account (top-right of the app).
  • Download at least one model you want to reach remotely. A 7B–14B model is the sweet spot for phone use — fast responses, modest VRAM, and smart enough for real work. (If you've got a 24GB+ card or a big unified-memory Mac, nothing stops you from serving a 70B.)
2. Get LM Link access:
  • Open LM Studio's LM Link page. Since the waitlist was removed, you can enable it directly.
3. On your iPhone or iPad:
  • Install the Locally app from the App Store.
  • Sign in with the same LM Studio account.
  • Your desktop shows up as an available host. Pick it, pick a model, and start chatting.

That's it. The encrypted mesh is established automatically by the shared account login.

What About Android?

Locally is iOS/iPadOS-only at launch, but Android users aren't locked out — you just take the manual road, which has existed for a while:

1. In LM Studio, open the Developer / Server tab and start the local server (default port 1234, giving you an OpenAI-compatible endpoint at http://localhost:1234/v1). From the CLI it's lms server start --port 1234. 2. Turn on "Serve on Local Network" (binds to 0.0.0.0 instead of loopback) so other devices can reach it. 3. Point any OpenAI-compatible Android chat client (Chatbox, LMSA, and others) at http://:1234/v1.

The big caveat: that exposes an unauthenticated server on your LAN. Only do it on a network you trust, and if you want access from outside your home, install Tailscale yourself and connect over the tailnet rather than opening a port to the internet. That's exactly the security model LM Link productizes — Android users just assemble it by hand. Full walkthrough in our LM Studio on Your Phone guide.

The Use Cases: Why You'd Actually Want This

Remote-access-to-your-own-rig is a different value proposition than tiny on-device models. Here's where it shines:

Run models your phone could never touch. This is the headline. A phone tops out around a 4B model on-device. With LM Link, you're driving whatever your desktop can run — 32B, 70B, big MoE models. Frontier-class local intelligence, from the couch. Private AI without a cloud subscription. Your data stays inside your own devices, inference happens on hardware you own, and there's no per-token API bill and no provider reading your conversations. For anyone handling sensitive material — legal, medical, financial, proprietary code — this is the responsible version of "AI on my phone." Long-context document and code work on the go. Phones can't hold a 128K-token context in memory, but your desktop can. Summarize a long contract, review a large diff, or query a big codebase from your phone while the heavy lifting happens at home. Use the GPU you're paying for, all the time. That RTX 4090 or M4 Max sits idle most of the day. LM Link turns it into an always-available personal inference server you can reach from anywhere — getting far more value out of hardware you already bought. A genuinely private voice/quick-question assistant. Fire off a quick question from your phone and get an answer from a real 14B model in seconds, without it leaving your tailnet. Shared family or small-team rig. One capable machine at home or in the office can serve several people's phones — a self-hosted, private alternative to handing everyone a ChatGPT subscription. Travel and remote work. As long as both devices can reach the tailnet, your home AI follows you. No data plan limits on model size, because the model isn't on the phone.

What It Is Not

A couple of honest caveats so expectations are right:

  • It's not offline. Unlike on-device models, LM Link needs a connection between your phone and the host. If your desktop is asleep or off the network, there's nothing to talk to. (Keep the host awake, or wake-on-LAN it.)
  • It's not local-to-the-phone privacy in the airplane-mode sense. It's private — encrypted, peer-to-peer, no cloud provider — but data does travel between your two devices. For most people that distinction doesn't matter; for true air-gapped use, on-device models are still the answer.
  • iOS first. Android is manual-config for now.

The Bigger Picture

LM Link quietly resolves the central tension of phone AI. You no longer have to choose between "small model that fits on my phone" and "powerful model in someone else's cloud." You can have a powerful model, on your own hardware, reachable from your phone, with the privacy properties of local AI intact.

That's the version of pocket AI that's actually worth using — and it's available today.


Ready to set it up? Follow the step-by-step LM Studio on Your Phone (LM Link) guide → Curious which model to serve? Browse the Model Library → and check fit with the Hardware Analyzer →