How They Fixed AI's Biggest Local Hardware Problem (with FreeToken)
FreeToken is an open-source inference engine that lets Mixture-of-Experts AI models run significantly faster than Ollama — but only once a model no longer fits comfortably inside a GPU's VRAM. This video breaks down how the tool works under the hood, from its expert caching to its bandwidth-adaptive CPU/GPU split, then puts it head-to-head against Ollama on a real coding task on an RTX 5090. The results show exactly where FreeToken wins, where Ollama still comes out ahead, and what its new desktop app can (and can't) do yet. 🔗 Relevant Links FreeToken: https://github.com/FlashML-org/FreeToken ❤️ More about us Radically better observability stack: https://betterstack.com/ Written tutorials: https://betterstack.com/community/ Example projects: https://github.com/BetterStackHQ 📱 Socials Twitter: https://twitter.com/betterstackhq Instagram: https://www.instagram.com/betterstackhq/ TikTok: https://www.tiktok.com/@betterstack LinkedIn: https://www.linkedin.com/company/betterstack 📌 Chapters: 0:00 Intro: The 3x Speed Claim 0:32 What Is FreeToken? 0:51 The Core Problem With MoE Models 1:56 The Fix: GPU Memory as a Cache 2:35 Prefill & Double Buffering 3:12 Decode Misses & the Q* Policy 4:17 The FTW Weight Format 4:35 Demo Setup: RTX 5090 4:56 Ollama vs FreeToken: How Each Handles Overflow 5:46 Real Coding Task: Ollama vs FreeToken 6:26 Live GPU Cache Resizing 7:05 The Catch: When Ollama Wins Instead 7:51 FreeToken's Desktop App 8:35 Final Thoughts 9:01 Outro