Rates verified 2026-07 · refreshed quarterly

Cloud GPUs & LLM APIs

Three ways to run a model you can’t run at home: rent a GPU by the hour, use a free API tier to prototype, or pay per token. Start with the one that matches your workload.

Cloud AI directory

Curated by Jakub Rusinowski · Entries verified 2026-07-08

This site is about running AI on your own hardware — but the cloud is the right tool surprisingly often: free API tiers for prototyping, rented GPUs for fine-tuning runs and try-before-you-buy, hosted open models when you want Llama-class output without owning a card. Here’s the honest map of what’s out there, and the math for when local wins.

Pricing notes are coarse by design (tiers change often); always confirm on the provider’s page. Most links are plain links; RunPod and Vast.ai are affiliate links.

Free LLM API tiers

Genuinely free API access — rate-limited, but real. Enough for prototyping, learning and low-volume tools before you spend anything on hardware or tokens.

NVIDIA NIM (build.nvidia.com)Free: Free hosted endpoints for Llama, Nemotron, Mistral & more with generous dev credits — sign-up only, no cardFree API credits for development; self-host NIM containers with NVIDIA AI Enterprise in productionTrying big open models (Llama 3.3 70B, Nemotron) over an OpenAI-compatible API before buying hardwareGoogle AI StudioFree: Gemini Flash free tier — the most generous standing free allowance of the big threeFree-of-charge tier with per-minute/per-day rate limits; paid tier unlocks higher limitsPrototyping against a frontier-family model at zero costGroqFree: Free developer tier with daily rate limits on open models (Llama, Qwen, Gemma)Pay-per-token on LPU hardware; famous for extreme tokens/secLatency-critical demos — hundreds of tokens/sec on 70B-class open modelsOpenRouterFree: Rotating set of :free model variants with daily request capsOne API key, 300+ models, pass-through pricing per modelComparing many models through one OpenAI-compatible endpoint before committingMistral La PlateformeFree: Free experiment tier with rate limitsPer-token pricing on Mistral models; EU-hostedEU-hosted API access with a free on-ramp — relevant for GDPR-sensitive prototypingCerebras InferenceFree: Free developer tier with daily token capsPer-token pricing on wafer-scale hardware; very high tokens/secFastest hosted open-model inference alongside Groq — good for agent loopsGitHub ModelsFree: Free playground + API rate limits for GPT-4o, Llama, Phi, Mistral with any GitHub accountIncluded with a GitHub account; production use routes to Azure AIZero-signup-friction experimentation if you already live on GitHubCloudflare Workers AIFree: Daily free allocation of neurons on every Cloudflare accountPer-neuron (unit) pricing beyond the free allocation; runs on Cloudflare edge GPUsSprinkling small-model inference (Llama 8B class) into edge apps without managing serversHugging Face Inference ProvidersFree: Small monthly free credit allowance on every HF accountRoutes to partner providers at pass-through prices; monthly included credits on free & PRO plansCalling almost any open-weight model on the Hub without picking a vendor first

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Local vs cloud: which should you use?

Editorial comparisons with a shared break-even calculator — honest about where each side wins. Cloud prices verified 2026-03-01.

Start hereLocal AI vs Cloud API Costs: The Real Break-Even Math (2026)Local AI vs cloud API costs compared with real numbers: hardware, electricity, and per-token prices. Break-even calculator shows when a GPU pays for itself.

Frequently asked questions

Where can I use AI models for free via API?

The most useful standing free tiers in 2026: NVIDIA NIM (build.nvidia.com) gives free hosted endpoints for Llama, Nemotron and Mistral-class models; Google AI Studio offers a generous Gemini Flash free tier; Groq and Cerebras run free rate-limited developer tiers on open models at extreme speed; OpenRouter exposes rotating free model variants; and GitHub Models works with any GitHub account.

What is the cheapest way to rent a GPU for AI?

Marketplace providers are cheapest: Vast.ai and TensorDock list consumer GPUs (RTX 3090/4090 class) from roughly $0.20–0.40/hour, RunPod adds per-second serverless billing, and Lambda/Hyperstack cover datacenter A100/H100 cards. For free experimentation, Google Colab still offers T4 notebook sessions at no cost.

Should I rent a cloud GPU or buy my own?

Rent when usage is occasional (fine-tuning runs, trying a 70B model before committing) — a few dollars per experiment. Buy when usage is sustained: at a few hours of daily inference, a $1,000–2,600 build typically undercuts rental within months. Our local-vs-cloud comparison pages include a break-even calculator.

What is the difference between hosted open-model APIs and frontier APIs?

Hosted open-model providers (Together, Fireworks, DeepInfra, Groq) serve open-weight models like Llama and Qwen per token — often 10–50× cheaper than frontier APIs, and you can migrate the same model to your own hardware later. Frontier APIs (OpenAI, Anthropic, Google) serve closed models with the highest quality ceiling.