192ms vs 1,006ms in a Voice Agent

LiveKit
3,738 views July 29, 2026

We benchmarked time to first token across the LLMs people actually put in voice agents. Gemma 4 31B on LiveKit Inference came back at 192ms. GPT-4.1 came back at 1,006ms. It's not that Gemma 4 is fast on its own. The same Gemma 4 31B served on other providers measures 1,876ms to first token on the same tests. 192ms is the result of a latency-optimized serving stack. #livekit #ai #voiceai

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close