192ms vs 1,006ms in a Voice Agent
We benchmarked time to first token across the LLMs people actually put in voice agents. Gemma 4 31B on LiveKit Inference came back at 192ms. GPT-4.1 came back at 1,006ms. It's not that Gemma 4 is fast on its own. The same Gemma 4 31B served on other providers measures 1,876ms to first token on the same tests. 192ms is the result of a latency-optimized serving stack. #livekit #ai #voiceai