Why Every Voice Agent Should Use Gemma 4 Now

LiveKit
125,509 views July 28, 2026

Gemma 4 31B on LiveKit Inference reaches 192ms time to first token, versus 911ms for Gemini 2.5 Flash and 1,006ms for GPT-4.1 in LiveKit's comparison. This video opens on that gap, then gets straight to how to use it: the one-line swap to put Gemma 4 in your own voice agent, a live console demo, and the numbers that matter. Sign up for LiveKit Cloud: https://cloud.livekit.io/signup?utm_source=youtube&utm_medium=video&utm_campaign=devrel&utm_content=RUFexZ8cZbs You will also see where the speed comes from (LiveKit's serving stack, not the model alone), how Gemma 4 stacks up on quality (88% task completion on the hotel receptionist eval, 75.6% on IFBench, 76.9% on tau-squared in LiveKit's benchmark run), and what it costs ($0.40 per 1M input tokens). 📚 Resources 📚 Agent docs: https://docs.livekit.io/agents/?utm_source=youtube&utm_medium=video&utm_campaign=devrel&utm_content=RUFexZ8cZbs 🤝 Join the Community: https://community.livekit.io/?utm_source=youtube&utm_medium=video&utm_campaign=devrel&utm_content=RUFexZ8cZbs #livekit #ai #voiceai

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close