~/blog --tail · 56 posts

The log

Benchmarks, model showdowns, and hard-won notes from running AI locally.

Looking for research? Read the Local AI Report →
All#Getting Started#Comparison#Benchmarks#GPU#Privacy#Fine-Tuning#Mobile#Cost
Oct 6, 2026Kolibri 1 and the European take on open-weight models: what Aleph Alpha shipped, and what it takes to run itModels · Kolibri 1 is Aleph Alpha's Apache-2.0 German-English 78B MoE. What the licence covers, how to run it, what the benchmarks show, and what 'European' means here.18 minSep 30, 2026Jev and the decision-model wave, explained — and how to run one on your own machineModels · TypeSafe's Jev answers typed questions with probabilities instead of text. How decision models work, what the benchmarks really say, and which open ones you can run locally.14 minSep 8, 2026Nex-AGI: Six Agentic Models, One of Which Fits Your GPUModels · Nex-AGI doesn't pretrain models — it post-trains other labs' open weights for agentic work, and six of the results are now in the library. Built on Qwen3.5 and DeepSeek bases, they run from 35B to 1.6 trillion parameters. Here's what each one is, where its numbers come from, and why exactly one of the six will load on a card you can buy.11 minSep 5, 2026The New Renaissance: What the Printing Press Actually Did, and What Open Models Are Doing NowModels · The printing press did not do what most people think it did. The Church was an early customer, not a panicked censor, and printing did not stay free. Read accurately, the fifteenth century says something precise about open-weight models — about why the narrow specialist is a 250-year-old industrial invention rather than a medieval one, and about which half of your craft just became free.13 minAug 26, 2026UNESCO's Nine Approaches to Governing AI — And What Comes After the TaxonomyPrivacy & Business · UNESCO's 2026 policy brief maps nine approaches lawmakers are using to regulate AI, from principles to liability, illustrated with real laws from Brazil to South Korea. The taxonomy is useful — but none of the nine, by itself, guarantees an organisation can demonstrate what actually happened across the AI lifecycle. Here is the map, read against the primary source, plus the gap it leaves and the operational architecture that closes it: Policy to Risk to Decision to Action to Evidence to Accountability, and Governance to Runtime to Usage to Consumption to Cost.18 minAug 25, 2026Kimi Linear, Explained: The Attention Architecture That Beat Full AttentionModels · Moonshot AI's Kimi Linear is the first architecture to beat full attention under a controlled comparison — while cutting KV cache 75% and decoding 6.3x faster at 1M tokens. Here's how Kimi Delta Attention actually works, transcribed from the open-source kernel: the state-update equation, why per-channel gating was the unlock, where the 75% comes from arithmetically, and what the paper does not prove.25 minAug 17, 2026How to Install and Set Up DeepSeek Harness (dsh)Getting Started · DeepSeek Harness (dsh) is DeepSeek's new open-source, plugin-based agent harness. Here's how to install it — npx quickstart, full source build, web UI and headless configuration, and your first cordis.yml plugin — plus what to watch out for while it's still a developer preview.11 minAug 17, 2026DeepSeek Harness Explained: The Concept, and Its Real Pros and ConsModels · DeepSeek Harness is DeepSeek's open-source agent harness, built so every part — model, tools, sandbox, even the UI — is a swappable plugin. Here's how that architecture actually works, and an honest breakdown of where it helps, where it hurts, and whether to adopt it while it's still a developer preview.14 minJul 27, 2026PageIndex: The Vectorless, Reasoning-Based RAG Alternative (2026 Guide)Models · PageIndex throws out vector databases, embeddings, and chunking. Instead it builds a table-of-contents tree from your documents and lets an LLM reason its way to the right section — the way a human expert would. Here's how vectorless, reasoning-based RAG works, how it hit 98.7% on FinanceBench, when it beats classic RAG, and how to run it (even against a local model).22 minJul 27, 2026The Complete Ollama Bible: Install, Run & Master Local LLMs (2026)Getting Started · The complete guide to Ollama: install it, pick a model that fits your GPU, choose a quantization, script the OpenAI-compatible API, build custom Modelfiles, tune for speed, and fix the errors everyone hits — with infographics and a private-RAG capstone.35 minJul 27, 2026The Complete LM Studio Bible: The GUI Way to Run Local LLMs (2026)Getting Started · The complete guide to LM Studio: download and chat with local models through a friendly GUI, read the fit badges, pick GGUF vs MLX, master the GPU-offload slider, chat with your documents, and serve an OpenAI-compatible API.30 minJul 27, 2026The Complete Open WebUI Guide: Build Your Own Private ChatGPT (2026)Getting Started · The complete guide to Open WebUI: turn a local model into a private, multi-user ChatGPT with document RAG, web search, image generation, voice, custom assistants, and safe remote access — all self-hosted.32 minJul 27, 2026The Complete Qwen Guide: Every Model Explained (Qwen 3 & 3.5, 2026)Models · The complete guide to Alibaba's Qwen: every model size explained across Qwen 3 and 3.5, thinking mode, MoE and Gated DeltaNet, native vision, 201 languages, which one fits your GPU, and how it compares to DeepSeek, Gemma and Llama.30 minJul 27, 2026The Complete RTX 5090 Guide for Local AI (2026)Hardware · The complete RTX 5090 guide for local AI: what its 32 GB and 1,792 GB/s of bandwidth deliver, which models it runs and how fast, how it beats the RTX 4090, power and cooling needs, dual-GPU builds, and whether it's actually worth it.33 minJul 27, 2026The Complete DeepSeek Guide: R1, V3.2 & V4 Explained (2026)Models · The complete guide to DeepSeek: how reasoning models work, the R1 distills you can actually run (8B/14B/32B), the giant V3.2 and V4 models, MIT licensing, privacy, and exactly when to reach for it — with infographics and a reasoning-workstation capstone.32 minJul 26, 2026The Digital Omnibus on AI: What Just Changed in the EU AI ActPrivacy & Business · The EU just rewrote the AI Act's timetable. Regulation (EU) 2026/1744 — the Digital Omnibus on AI — delays the heaviest high-risk obligations to 2027 and 2028, ties them to a readiness trigger instead of a fixed date, and adds two new prohibitions that bite in 2026. Here's what changed, in plain English, with infographics.9 minJul 22, 2026Rent a GPU by the Hour: When Vast.ai and RunPod Beat BuyingHardware · You checked whether your GPU can run that 70B model and it can't. Before you drop $1,600 on a new card, here's when renting a cloud GPU by the hour on Vast.ai or RunPod is the smarter move — with the exact break-even math and step-by-step guides.8 minJul 21, 2026NVIDIA Cosmos 3: The Open World Model That Teaches Robots to Think Before They ActModels · NVIDIA's Cosmos 3 is an open, omnimodal 'World Foundation Model' for Physical AI — one system that understands and generates text, images, video, ambient audio and robot actions. Here's what a world model actually does, how its two-tower architecture 'thinks before it acts,' the Super / Nano / Edge sizes, what hardware it needs, and why NVIDIA gave it away — with infographics you can skim.10 minJul 16, 2026Thinking Machines' Inkling: A Trillion-Parameter Open Model That Sees and HearsModels · Mira Murati's Thinking Machines just shipped Inkling — a 975-billion-parameter, Apache-2.0 model that natively understands text, images and audio, handles a 1-million-token context, and fires only ~41B parameters per token. Here's what it is, how it benchmarks, what it takes to run, and the open-weights bet behind giving it away — with infographics you can skim.9 minJul 15, 2026vLLM in 2026: The V1 Engine, Compatible APIs, and When to Actually Use ItHardware · vLLM's V1 engine is now the default, its server speaks the OpenAI and Anthropic APIs out of the box, and independent 2026 benchmarks put it 2–29× ahead of Ollama under concurrent load. Here's what changed this year, how it works in plain terms, and exactly when it's the right tool — and when Ollama or LM Studio still win.8 minJul 15, 2026Bonsai 27B: The Largest AI Model That Runs on Your PhoneMobile · PrismML just fit a full 27-billion-parameter multimodal model into 3.9 GB — small enough to run on an iPhone. Here's how the 1-bit and ternary builds work, how much intelligence survives the compression, how fast it actually runs, and which build you should use.9 minJul 14, 2026The EU AI Act Compliance Guide: Timeline, Who's Affected & What to DoPrivacy & Business · The EU AI Act is here — and the 2026 Digital Omnibus just moved the deadlines. A plain-English guide to the compliance timeline, who's actually affected, the four risk tiers, the fines, and a 7-step action plan — with infographics you can skim.14 minJul 13, 2026The Enterprise Shift to Local LLMs: Data Sovereignty, Token Economics, and Models Fine-Tuned for Your BusinessFine-Tuning · Production AI inference is leaving the public cloud — public-cloud share fell from 56% to 41% in one year. The three forces driving enterprises to local LLMs: data sovereignty, token economics, and fine-tuning models on their own data.11 minJul 9, 202647 Open Datasets for Fine-Tuning LLMs in 2026 — and How to Use Them in One HourFine-Tuning · Fine-tuning went from research lab to weekend project — a 7B QLoRA run fits on a 6GB gaming GPU or a 16GB MacBook. The 47 open datasets worth knowing in 2026.6 minJul 8, 2026You Don't Need Hardware to Use Powerful Models for FreeGetting Started · No GPU, no problem. Free API tiers from NVIDIA, Google, Groq, and others give you real access to powerful models — if you understand the rate limits, privacy trade-offs, and fine print.8 minJun 27, 2026Your Home GPU in Your Pocket: Run LM Studio Models From Your Phone With LM LinkMobile · LM Studio's new LM Link and the Locally iPhone app let you run your biggest local models — even 70B+ — from your phone. The model stays on your desktop; only the chat travels, end-to-end encrypted. Here's how it works and why it matters.9 minMay 29, 2026Do You Need an Expensive PC to Run Local AI? The Honest AnswerHardware · Most articles about local AI assume you have a high-end gaming PC. The reality is friendlier than that. Here's exactly what hardware you need — and what you can already run right now.7 minMay 29, 2026Your First Local AI: Get Up and Running in 15 MinutesGetting Started · No technical background needed. This step-by-step guide walks you through installing your first local AI, choosing your first model, and having your first conversation — all in about 15 minutes.8 minMay 28, 2026What Is a Local LLM? A Plain-English Guide for Complete BeginnersGetting Started · You've heard about ChatGPT and AI assistants — but what exactly is a 'local LLM', and what does it mean to run AI on your own computer? This guide explains it all without the jargon.7 minMay 28, 2026ChatGPT vs. Local AI: The Complete Honest ComparisonGetting Started · ChatGPT and Claude are convenient and powerful. Local AI is free, private, and yours. This honest comparison covers cost, privacy, quality, and setup so you can decide which one actually fits your life.8 minMay 27, 2026Multi-Agent LLM Systems: How AI Orchestrates Itself to Solve Complex ProblemsModels · A single LLM prompt has limits. Multi-agent systems — where LLMs plan, delegate, and verify each other's work — can tackle problems that would be impossible in a single pass. Here's how they work and what they can actually do.11 minMay 23, 2026LLM Quantization Explained: How a 70B Model Shrinks from 140 GB to 8 GBModels · A raw 70B parameter model needs 140 GB of storage and VRAM. Quantization gets it down to 8 GB with surprisingly small quality loss. Here's how it works and which level to choose.9 minMay 20, 2026LLM Context Windows Explained: Why 1 Million Tokens Changes EverythingModels · Context windows have grown from 4,096 tokens in 2022 to over 1 million in 2026. This isn't an incremental improvement — it fundamentally changes what AI can do with a single prompt.10 minMay 4, 2026Gemma 4 27B: GPT-4 Level AI on Your Gaming GPUModels · Google DeepMind's Gemma 4 27B fits in 14 GB VRAM and hits 85 tokens/second on an RTX 4090 — delivering frontier-class intelligence without a data center.8 minMay 4, 2026DeepSeek V4 Pro: How a 1.6 Trillion Parameter Model Beat EveryoneModels · DeepSeek V4 Pro has taken the #1 spot on every major open-weight benchmark. Here's the architecture behind it, what it costs to run, and why it matters for the future of open AI.9 minMay 4, 2026Kimi K2.6: The Model Built for Autonomous Coding AgentsModels · Moonshot AI's Kimi K2 series is purpose-built for agentic AI — models that don't just answer questions but plan, use tools, write code, and debug it autonomously. Here's what makes it different.7 minMay 4, 2026Qwen 3.5: When Text, Vision, and Video Understand Each OtherModels · Alibaba's Qwen 3.5 (397B) is the first truly native multimodal open-weight model — not a language model with a vision plugin, but a single architecture that thinks in text, images, and video simultaneously.8 minMay 4, 2026GLM-5 / GLM-5.1: Why MIT License Matters for AI in ProductionModels · Zhipu AI (Z.ai)'s GLM-5 series is one of the most capable models available under a true MIT license — meaning no restrictions, no royalties, no legal gray areas. Here's what you can build with it.7 minApr 13, 2026Local LLMs for Privacy-Safe Document Analysis — A Practical GuidePrivacy & Business · Your contracts, financial records, and medical documents should never be processed by cloud AI. Here's how to set up a local document analysis stack that keeps sensitive information entirely on your hardware.9 minApr 12, 2026Best Small LLMs for Low-End Hardware — Running AI on 4 GB VRAMHardware · You don't need a $1,500 GPU to run a useful local AI. These models run fast and smart on integrated graphics, old gaming GPUs, and even a decent laptop. Here's the 2026 guide to small but capable.7 minApr 12, 2026Why Companies Fine-Tune Open-Source LLMs Instead of Paying Per-Token API BillsPrivacy & Business · A fine-tuned 7B model can outperform GPT-4 on specific tasks at 1/200th the running cost. Here's the full business case, and how companies build defensible AI moats with open-source fine-tuning.11 minApr 11, 2026DeepSeek R1 vs GPT-4o — Can a Free Local Model Match the Cloud?Models · DeepSeek R1 runs on your own hardware and costs nothing per query. GPT-4o costs $15 per million tokens and sends your data to OpenAI's servers. But which one actually wins on quality?8 minApr 11, 202612 Real-World Use Cases for Running a Local LLM in 2026Getting Started · From air-gapped legal review to offline coding assistants and on-device customer support — local LLMs have become production-ready tools. Here are the best use cases with setup tips.10 minApr 10, 2026How to Run LLMs Completely Offline — The Air-Gap GuidePrivacy & Business · No internet, no cloud dependency, no data leaving your machine — ever. Whether it's for security, compliance, or survival, here's how to set up a fully air-gapped local AI that works anywhere.8 minApr 10, 2026Gemma 4 Review: Google's Best Open-Source Model — All Sizes ComparedModels · Google's Gemma 4 landed in April 2026 with a 92.4% MMLU score, native multimodal vision, audio support, and a 1B model that runs on a phone. Here's the full breakdown.9 minApr 9, 2026Open WebUI vs LM Studio vs Jan — Which Local AI Interface Is Right for You?Models · Three great tools, three different use cases. Whether you want a polished desktop app, a self-hosted team platform, or a lightweight offline interface, here's how to choose.7 minApr 8, 2026Llama 4 Review — The Best Open-Source LLM Yet?Models · Meta's Llama 4 Scout and Maverick mark a major leap for open-source AI: MoE architecture, 10 million token context, and multimodal capabilities. Here's what they actually deliver in practice.8 minApr 7, 2026How to Set Up a Private AI Assistant for Your Business in 2026Privacy & Business · Your employees are already using ChatGPT with company data. Here's how to give them a better AI tool that keeps every conversation on your own servers — for less than $50/month.9 minApr 5, 2026Best GPU for Local AI Under $500 in 2026 (Tested & Ranked)Hardware · You don't need a $1,600 RTX 4090 to run impressive local AI. These five GPUs under $500 cover every budget from $200 to $500 — with real benchmark numbers.7 minApr 3, 2026The State of Local AI in 2026: What's Changed in 12 MonthsModels · A year ago, running a 70B model required a $10,000 server. Today you can do it on a MacBook Pro. Here's everything that changed in local AI over the past 12 months.9 minApr 1, 2026RTX 5090 vs RTX 4090 for Local LLMs: Is It Worth $2,000?Hardware · The RTX 5090 delivers 213 tokens/sec versus the 4090's 165 — a 29% speed boost with 33% more VRAM. Here's who should upgrade and who should wait.8 minMar 28, 2026How to Start Fine-Tuning Your Local LLM: A Beginner's GuideFine-Tuning · Fine-tuning turns a general-purpose AI into one that knows your business, your writing style, and your domain. Here's how to start — even if you've never trained a model before.10 minMar 26, 2026Who Is Winning the Local LLM Race in 2026?Models · Qwen 3.5 topped reasoning benchmarks, Llama 4 brought MoE to the masses, Kimi K2.5 hit HumanEval 99. Who actually wins when you run them on your own hardware?9 minMar 22, 2026Can You Run a Local LLM on Your Phone? Yes — and It's Better Than You ThinkMobile · Modern smartphones are powerful enough to run real AI models completely offline. No internet. No subscription. No privacy concerns. Here's how to set it up today.6 minMar 18, 2026How Companies Are Saving Thousands by Running AI LocallyPrivacy & Business · From legal compliance to slashed API bills, businesses running their own LLMs are gaining a real edge. Here's the full business case for private AI in 2026.8 minMar 15, 2026Why Running a Local LLM Is Cheaper and More Secure Than You ThinkPrivacy & Business · Cloud AI bills add up fast. Running your own LLM locally can cost as little as $2/month in electricity — while keeping every word you type off someone else's server.7 min