Cloud GPUs & LLM APIs
Three ways to run a model you can’t run at home: rent a GPU by the hour, use a free API tier to prototype, or pay per token. Start with the one that matches your workload.
Cloud AI directory
Curated by Jakub Rusinowski · Entries verified 2026-07-08
This site is about running AI on your own hardware — but the cloud is the right tool surprisingly often: free API tiers for prototyping, rented GPUs for fine-tuning runs and try-before-you-buy, hosted open models when you want Llama-class output without owning a card. Here’s the honest map of what’s out there, and the math for when local wins.
Pricing notes are coarse by design (tiers change often); always confirm on the provider’s page. Most links are plain links; RunPod and Vast.ai are affiliate links.
Free LLM API tiers
Genuinely free API access — rate-limited, but real. Enough for prototyping, learning and low-volume tools before you spend anything on hardware or tokens.
Cloud GPU rental
Rent a GPU by the hour instead of buying one. The right answer for fine-tuning runs, batch jobs and “try a 70B before committing to a $2,000 build”.
Hosted open-model APIs
Open-weight models (Llama, Qwen, Mistral, DeepSeek) served per token by specialists — the middle ground between frontier APIs and your own hardware.
Frontier model APIs
The closed frontier models. Highest quality, zero setup, per-token billing — the benchmark every local setup gets compared against.
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Local vs cloud: which should you use?
Editorial comparisons with a shared break-even calculator — honest about where each side wins. Cloud prices verified 2026-03-01.
Start hereLocal AI vs Cloud API Costs: The Real Break-Even Math (2026)Local AI vs cloud API costs compared with real numbers: hardware, electricity, and per-token prices. Break-even calculator shows when a GPU pays for itself.Frequently asked questions
Where can I use AI models for free via API?
The most useful standing free tiers in 2026: NVIDIA NIM (build.nvidia.com) gives free hosted endpoints for Llama, Nemotron and Mistral-class models; Google AI Studio offers a generous Gemini Flash free tier; Groq and Cerebras run free rate-limited developer tiers on open models at extreme speed; OpenRouter exposes rotating free model variants; and GitHub Models works with any GitHub account.
What is the cheapest way to rent a GPU for AI?
Marketplace providers are cheapest: Vast.ai and TensorDock list consumer GPUs (RTX 3090/4090 class) from roughly $0.20–0.40/hour, RunPod adds per-second serverless billing, and Lambda/Hyperstack cover datacenter A100/H100 cards. For free experimentation, Google Colab still offers T4 notebook sessions at no cost.
Should I rent a cloud GPU or buy my own?
Rent when usage is occasional (fine-tuning runs, trying a 70B model before committing) — a few dollars per experiment. Buy when usage is sustained: at a few hours of daily inference, a $1,000–2,600 build typically undercuts rental within months. Our local-vs-cloud comparison pages include a break-even calculator.
What is the difference between hosted open-model APIs and frontier APIs?
Hosted open-model providers (Together, Fireworks, DeepInfra, Groq) serve open-weight models like Llama and Qwen per token — often 10–50× cheaper than frontier APIs, and you can migrate the same model to your own hardware later. Frontier APIs (OpenAI, Anthropic, Google) serve closed models with the highest quality ceiling.
Keep going
By budget
By use case
- Local vs Cloud AI for Coding: Copilot-Class Help Without the Meter?
- Local vs Cloud AI for Writing: Drafts, Editing, and the Privacy of an Unsent Sentence
- Local vs Cloud AI for RAG: Where Your Documents Live Decides
- Local vs Cloud AI for Agents: Token Furnaces Meet the Meter
- Local vs Cloud AI for Chatbots: Self-Hosting the Front Desk
By team size
By privacy & compliance
Prefer to own the hardware?