Local models
Running models on your own hardware: Ollama, llama.cpp, open weights and on-device AI
Latest stories See all →
- AI for IoT: What Is Edge AI, and What Will It Enable?
- How to Choose the Best AI PC for Your Small Business
- AMD Brings Local AI to Autodesk University 2026
- Can Local LLMs vs GPT-4 Handle Business Logic?
- Balena Appoints Leon Hurst as CEO as It Scales Platform for IoT and Edge AI
- webAI Frontline Cuts Aircraft Manual Search by 66%, Offline, on a Single iPadCompany News/
- Build Qt Edge AI applications faster on Dragonwing IQ‑9075 with Qt Onboard
- We had a frontier AI coach a small local model through Camel. It found 99 things wrong for humans too
- Cloud vs. Local AI Agents: Where Should Yours Run?
- A Brain Too Big to Carry — On-Device vs Datacenter Inference
- AnnouncementsCline Desktop: An open-source app for open-weight modelsSeptember 14, 2026 • Etisha GargMeet Cline Desktop, an open-source app for open-weight models. Run parallel AI agents, automate recurring tasks, choose from 300+ models, and extend your workflow with plugins, MCP servers, and skills.
- Chinese AI models dominate OpenRouter’s US token consumption. It can now guarantee that traffic stays entirely in the US.
- Portable Intelligence: Why AI Belongs Where Work Happens
- Introducing Dependency Firewall with On-Device Supply Chain Protection
- Introducing Dependency Firewall with On-Device Protection
- Kimi K3 Runs Locally at 8,192 Tokens — and Colibrì Ignores the CTX You Set
- Speculative Decoding in Llama.cpp: How to Actually Speed Up Local LLMs
- Could "Pacing the Frontier" Rhetoric Lead to a Ban on Local AI?
- Garry Tan wants US open-weight AI labs to 'distill' frontier models, too
- “Machine translation is still broken for most of the world’s languages”: Cohere builds non-reasoning for a reason
- Latency Optimization Levers for Open-Weight LLM inference:Part-2
- The Local AI Hardware Trap: Why VRAM, Bandwidth, and Resale Value Matter More Than Benchmarks
- NeoHorse-1-4B: How to Run This Self-Improving 4B Model Locally
- Agnes-3.0-Flash Preview: The Open-Weight Model, Explained
Videos
- Obscura + Ollama Local AI Web Scraping with a Rust Headless Browser
- Prime Agent + Ollama and LM Studio: An Agent That Spawns Its Own Agents
- Vector embeddings in ClickHouse with Ollama and OpenAI
- Run Claude Code with Ollama for 99% Cheaper AI
- OpenLumara + Ollama: Local Agent Trying to Beat OpenClaw and Hermes
- Local GenAI on Jetson: OSS models using different inferencing frameworks: Ollama, llama.cpp, & vLLM
- This Open-Source Tool Replaces Ollama + LangChain + Your UI
- NullClaw + Ollama Local Setup with Telegram
- ZeroClaw + Ollama: The Fastest OpenClaw Fork Yet? - Local Setup & Review
- PicoClaw + Ollama + Telegram: Run AI Agent on $10 Hardware - Full Local Setup Guide
- 🤩 How to Use Claude Code Free with Ollama
- Open NLP Meetup 15: Web Data Extraction with Trafilatura & Deploying with Hayhooks and Open WebUI
Podcasts
Recent episodes
- Firmware Analysis, Linux Malware, Future of AI - BTS #82
- EP 192: Talking Apple Silicon with Apple's Tom Boger, Kaiann Drance, and Sri Santhanam
- Elon Musk’s Event Almost Nobody Was Invited To | The Brainstorm 148
- We unfolded the iPhone Duo
- How Open-Source is Reshaping the AI Infrastructure Stack
- SED News: The NVIDIA-Hugging Face Deal, China’s Proxy Economy, the Open Weight Surge
- SED News: The NVIDIA-Hugging Face Deal, China’s Proxy Economy, the Open Weight Surge
- SED News: The NVIDIA-Hugging Face Deal, China’s Proxy Economy, the Open Weight Surge
- SED News: The NVIDIA-Hugging Face Deal, China’s Proxy Economy, the Open Weight Surge
- Aaron Levie on Why Open AI Wins
- Aaron Levie on Why Open AI Wins
- Aaron Levie on Why Open AI Wins