ai-PULSE 2025: Real-time AI: building 100x faster inference

Scaleway
21 views December 30, 2025

Inference speed is the bottleneck holding back the next generation of AI applications. Agentic workflows, real-time creation, and test-time compute scaling all demand much faster token generation at runtime. This talk reveals Kog's path to 100x faster AI inference through systematic GPU optimization and novel Transformer architectural innovations that circumvent hardware constraints. When AI operates at 10,000 or 100,000 tokens per second, computers stop being static rule-execution engines and become smart, real-time adaptive platforms. Programming in natural language becomes practical. Complex games and applications generate themselves dynamically. Agents and Deep Research are instant. Gaël, founder and CEO of Kog, will share the unique research insights and GPU engineering breakthroughs used to build the future of real-time AI computing With Gaël Delalleau Founder and CEO of Kog See more 👉 https://www.ai-pulse.eu/

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close