PyTorch Day India 2026 Why Speech Recognition Is Becoming an LLM Problem! Abhigyan Raman, Sarvam ai

PyTorch
242 views February 16, 2026

Problem Framing: Speech is the most natural interface for humans—but historically the hardest modality to scale. In multilingual regions like India, text-first AI systems struggle for deep penetration making way for speech as the default gateway to equitable AI access. Core Thesis: ASR has transitioned from acoustic-first pipelines to an LLM-first paradigm. Modern decoder-only Audio LLMs treat speech as another tokenized modality, unlocking better multilingual scaling, reasoning, and adaptation. What This Talk Covers: Evolution of ASR architectures: CTC → Encoder–Decoder (with cross-attention) → Decoder-only Audio LLMs Why alignment (audio → token space) is the key technical unlock Tradeoffs across pradigmns: latency, streaming, robustness, multilinguality Practical post-training strategies for domain- and language-specific ASR Why It Matters: This shift collapses the boundary between speech recognition and language understanding—making speech a first-class citizen in foundation model stacks.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close