Run LM Studio Models on Your Phone (LM Link) — 2026 Guide

By Jakub RusinowskiLast updated 10 min readBeginner

There are two ways to put AI on your phone. On-device inference runs a small 1B–4B model directly on the phone's chip — fully offline, but limited (see Running LLMs on Your Phone). Remote inference runs the model on your powerful desktop and streams the chat to your phone — so you can drive a 14B, 32B, or even 70B model from your pocket.

This guide covers the remote path with LM Studio's LM Link and the Locally app, plus a manual fallback for Android and advanced users.


How It Works

  • The model runs on your desktop (Mac, Windows, or Linux) inside LM Studio. The weights never leave that machine.
  • Your phone is a thin client: it sends prompts and streams back tokens.
  • The connection is end-to-end encrypted and peer-to-peer — your devices are never exposed to the public internet.

LM Studio splits this into two pieces:

PieceWhat it isWhere
LocallyThe chat appiPhone / iPad (App Store)
LM LinkThe encrypted remote-access layerBuilt into LM Studio + Locally

LM Link runs on top of a custom Tailscale mesh VPN. Both devices sign into the same LM Studio account, and that shared identity authenticates the link. Only a device-discovery list ever touches LM Studio's servers — your chats stay on your devices.

Status (June 2026): Requires LM Studio 0.4.16+. The early request-gated Preview waitlist was removed on June 8, 2026, so LM Link is open to everyone. It is free during the Preview; Locally is iPhone/iPad only at launch (no Android yet).

What You Need

  • A host machine (Mac, Windows, or Linux) with a capable GPU or Apple Silicon, running LM Studio 0.4.16 or later.
  • A free LM Studio account (sign in from the app's top-right).
  • At least one downloaded model. 7B–14B is the sweet spot for phone use — fast and capable. Bigger cards/Macs can serve 32B–70B.
  • An iPhone or iPad for the official path (or any OpenAI-compatible client for the manual Android path below).

Step 1 — Prepare the desktop

  1. Update LM Studio to 0.4.16+.
  2. Sign in to your LM Studio account.
  3. Download a model to serve. From the app press Cmd/Ctrl + Shift + M to search models, or use the CLI:

``bash lms get lmstudio-community/Qwen2.5-7B-Instruct-GGUF ``

Step 2 — Enable LM Link

  1. Open the LM Link page inside LM Studio.
  2. Turn it on (no waitlist required since the June 8 build).

Step 3 — Set up your phone

  1. Install Locally from the App Store.
  2. Sign in with the same LM Studio account.
  3. Your desktop appears as an available host — select it.

Step 4 — Chat

  1. Pick a model your desktop has loaded.
  2. Send a message. Inference runs on the desktop; the response streams to your phone over the encrypted link.

That's the whole setup. No IP addresses, no port forwarding, no QR codes — the shared account login establishes the mesh automatically.


Part 2 — The Manual Path: Local Server over LAN / Tailscale (Android & advanced)

Locally is iOS-only for now, but any OpenAI-compatible chat app can reach LM Studio's built-in server.

Step 1 — Start the LM Studio server

  • In LM Studio, open the Developer / Server tab and click Start Server, or from the CLI:

``bash lms server start --port 1234 ``

  • This exposes an OpenAI-compatible API at http://localhost:1234/v1 with endpoints like /v1/chat/completions, /v1/models, and /v1/embeddings.

Step 2 — Make it reachable from your phone

  • In the Server settings, enable "Serve on Local Network" (this binds to 0.0.0.0 instead of 127.0.0.1).
  • Find your desktop's LAN IP (e.g. 192.168.1.42).

Step 3 — Point a mobile client at it

  • Install an OpenAI-compatible Android client (e.g. Chatbox, LMSA) or, on iOS, an app that takes a custom base URL.
  • Set the base URL to http://<desktop-LAN-IP>:1234/v1 and leave the API key blank (or any placeholder).
  • Select the loaded model and chat.

Step 4 (recommended) — Secure remote access with Tailscale

  • The LAN bind has no authentication — only use it on a network you trust.
  • To reach your desktop from outside your home without opening a port to the internet, install Tailscale on both the desktop and phone, sign into the same tailnet, and use the desktop's Tailscale IP (100.x.y.z) instead of the LAN IP.
  • This is the same security model LM Link automates for you — you're just wiring it by hand.

Choosing a Model to Serve

Serving from a desktop sidesteps the phone's own memory limit entirely. If you would rather run the model on the device, note that iOS terminates an app that exceeds its memory budget and that model downloads stop when the app leaves the foreground.

Desktop VRAM / Unified RAMGood remote modelWhy
8 GBLlama 3.1 8B / Qwen2.5 7B (Q4)Fast, fits comfortably
12–16 GBQwen2.5 14B / Mistral Small (Q4)Big quality jump, still snappy
24 GBQwen 3 32B / DeepSeek R1 32B (Q4)Desktop-class reasoning on your phone
32 GB+ / big MacLlama 3.3 70B (Q4)Frontier-class local intelligence remotely

Use the Hardware Analyzer to confirm fit, and browse the Model Library — each model page has a one-click lms get command and an LM Link shortcut.


Security & Privacy Notes

  • LM Link: end-to-end encrypted over a private Tailscale mesh; devices never exposed to the public internet; chats stay on-device; inference on your hardware.
  • Manual LAN server: the default loopback bind (127.0.0.1) is what protects you. The moment you bind to 0.0.0.0, anyone on the network can hit the API — secure it at the network level or front it with an authenticating reverse proxy, and prefer Tailscale over port-forwarding for remote access.
  • Either way, no third-party cloud AI provider sees your prompts — that's the whole point.

Troubleshooting

  • Phone can't see the desktop (LM Link): confirm both devices are signed into the same LM Studio account and that the desktop is awake and online.
  • Connection refused (manual): the server isn't bound to the network — enable "Serve on Local Network" and check the desktop firewall allows port 1234.
  • Works on Wi-Fi, not on cellular: LAN IPs only work on the same network. Use Tailscale (or LM Link) for off-network access.
  • Slow responses: the desktop is doing the work — a smaller quant or a 7B–14B model improves latency over the link.
  • Desktop asleep: remote access needs the host online. Disable sleep, or use wake-on-LAN.

Use Cases

  • Run models far beyond phone limits — 32B/70B from your pocket, powered by your home GPU.
  • Private AI with no subscription — your data stays on your devices; no per-token API bills.
  • Long-context document & code work that won't fit in a phone's memory.
  • Get value from idle hardware — your GPU becomes an always-available personal inference server.
  • Family/team rig — one machine serves several people's phones privately.
  • Travel & remote work — your home AI follows you across the tailnet.

Remember: LM Link is remote, not offline. It needs a connection between phone and host. For true airplane-mode AI, use on-device apps — see Running LLMs on Your Phone.